Skip to content
Merged
6 changes: 5 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -74,8 +74,12 @@ PORT=8000
# AGENT_ENABLED=true
# AGENT_SERVICE_URL=http://localhost:8001
# AGENT_USER_ID_KEY=local-dev-key
# Upstream timeout in seconds, and the cap on one relayed chat stream.
# Upstream timeout in seconds, and the cap on one relayed chat stream outside
# the agents service's run queue.
# AGENT_REQUEST_TIMEOUT_S=190
# Hard cap in seconds on one chat turn, queue wait included; keep it below
# anyplot-api's Cloud Run --timeout (600).
# AGENT_TURN_MAX_S=590

# ============================================================================
# AI Services (optional)
Expand Down
19 changes: 13 additions & 6 deletions agents/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,16 +4,18 @@ This directory holds the anyplot agent network: the service that lets an admin p

## What is built

The runtime core runs locally: the agents, the plot pipeline, the guardrail plugins, the render layer and the private `/v1` service. Not built yet: the Cloud Run sandbox render backend (it waits for spike S), the container image, the deploy, the regression harness and the evals.
The runtime core runs locally: the agents, the plot pipeline, the guardrail plugins, the render layer, the private `/v1` service with its run queue, and the theme toggle. Not built yet: the Cloud Run sandbox render backend (it waits for spike S), the container image, the deploy, the regression harness and the evals.

The model is **Claude Haiku 5.5 on Vertex AI** (`claude-haiku-5-5`) by default. **Gemini 3.8 Flash** is the second arm: set `AGENT_PROVIDER=gemini` together with Gemini model ids, so the two can be compared on price and quality later. Every agent and the scope judge run on the configured provider.

## What lives here

| Path | Contents |
|---|---|
| `main.py` | The `anyplot-agents` FastAPI service: the `/v1` routes the BFF (`api/routers/agent.py`) calls, the caller check, in-memory session and artifact services, and the idle sweeper |
| `stream.py` | The `anyplot/1` stream translator: ADK events in, sanitised `ready`, `status`, `message`, `plot`, `refusal`, `error` and `done` events out |
| `main.py` | The `anyplot-agents` FastAPI service: the `/v1` routes the BFF (`api/routers/agent.py`) calls, the caller check, the run registry behind `409 run_active`, in-memory session and artifact services, and the idle sweeper |
| `stream.py` | The `anyplot/1` stream translator: ADK events in, sanitised `ready`, `status` (including `queued`), `message`, `plot`, `refusal`, `error` and `done` events out |
| `anyplot/run_queue.py` | The run queue in front of every `/messages` turn: one run in flight, one start a minute, a 600-second maximum wait, a `premium` lane that nothing sets yet |
| `anyplot/theme_render.py` | The theme toggle: renders another theme of a finished version from its stored run form, behind waiting pipeline renders, with no model call |
| `anyplot/agent.py` | The root agent `anyplot`, the `ALL_AGENTS` registry and `app` (the ADK `App` with its plugins), which `adk web` loads |
| `anyplot/models.py` | The only place that builds a model or a model client: `make_model`, `make_content_config` and `make_judge_client`, for Claude on Vertex AI and for Gemini |
| `anyplot/policy.py` | Composes each agent's static instruction from `anyplot/prompts/` and the catalogue's prompt sources, read verbatim; the fixed refusals; the data fences |
Expand All @@ -22,7 +24,7 @@ The model is **Claude Haiku 5.5 on Vertex AI** (`claude-haiku-5-5`) by default.
| `anyplot/sub_agents/` | One single-turn adapter per enabled library and the tool-less reviewer |
| `anyplot/tools/session.py` | The root's tools: `get_dataset_profile`, `get_spec_brief`, `get_current_code`, `set_bindings` and the `plot_pipeline` workflow |
| `anyplot/plugins/` | `ScopeGuardPlugin`, `BudgetPlugin`, `ToolSafetyPlugin` and the request ledger they share |
| `anyplot/render/` | The render contract, the probe harness, the host gates (R1-R3, advisory G3/G5/G7/G8), PNG hardening, the render store, and the `fake`, `local` and `sandbox` backends |
| `anyplot/render/` | The render contract, the probe harness, the host gates (R1-R3, advisory G3/G5/G7/G8), PNG hardening, the render store, the `fake`, `local` and `sandbox` backends, and `serial.py`, the one render semaphore in front of every backend |
| `anyplot/opening.py`, `session_state.py`, `services.py`, `briefs.py` | Opening a session and taking in a dataset, the server-set session state, the process-wide stores, and the spec and dataset briefs |
| `anyplot/dev_fixture.py` | The development-only session seed from an eval fixture case |
| `anyplot/settings.py` | `AgentSettings`, read from `AGENT_*` environment variables |
Expand Down Expand Up @@ -50,7 +52,10 @@ Three rules hold for everything here:
| `AGENT_LIBRARIES` | `matplotlib,seaborn` | Enabled libraries; each needs a phase-1 runtime (`matplotlib`, `seaborn`), others are refused at startup |
| `AGENT_RENDERER` | `sandbox` | `sandbox`, `local` (Docker, development only), `fake` (fixture PNGs, development and test only) or `remote` |
| `AGENT_RENDER_IMAGE` | `anyplot-agents:dev` | Image the `local` renderer runs |
| `AGENT_RENDER_CONCURRENCY` | `2` | Theme renders at the same time |
| `AGENT_RENDER_CONCURRENCY` | `1` | Theme renders at the same time, for every backend (`render/serial.py`); serial, because one 4 GiB instance holds one sandbox safely (spikes S and S2) |
| `AGENT_RUN_CONCURRENCY` | `1` | Pipeline runs (whole `/messages` turns) in flight; the run queue holds the rest |
| `AGENT_RUNS_PER_MINUTE` | `1` | Runs that may start within any 60 seconds (a sliding window) |
| `AGENT_QUEUE_MAX_WAIT_S` | `600` | Longest wait in the run queue, after which the run ends with `capacity`; the queue holds rate x wait / 60 entries (10) and answers `503 capacity` beyond that. Through the BFF a turn waits at most about 385 s until anyplot-api's request timeout is raised (`AGENT_TURN_MAX_S` in `docs/reference/api.md`) |
| `AGENT_MAX_LLM_CALLS` | `12` | LLM calls per request |
| `AGENT_REQUEST_TOKEN_BUDGET` | `80000` | Tokens per request |
| `AGENT_DAILY_TOKEN_BUDGET` | `1000000` | Tokens per user and day |
Expand Down Expand Up @@ -114,7 +119,7 @@ Three rules hold for everything here:

3. Open `http://localhost:8002`, choose `anyplot`, and send "Create the plot".

`adk web` and `adk api_server` are unauthenticated and let the client choose the user id, so run them only on your own machine. Every message costs model calls: a "Create plot" is about four (root twice, the adapter, the reviewer) plus one judge call for free text.
`adk web` and `adk api_server` are unauthenticated and let the client choose the user id, so run them only on your own machine. Every message costs model calls: a "Create plot" is about four (root twice, the adapter, the reviewer) plus one judge call for free text. Runs that `adk web` starts bypass the run queue, which sits in the `/v1` service: the rate and concurrency limits apply only to the service. Their renders are still serial, because every render goes through `Services.backend`, which puts the one render semaphore (`render/serial.py`) in front of the backend.

### Run the service

Expand All @@ -124,6 +129,8 @@ uv run uvicorn agents.main:app --port 8001

The service needs the header `X-Anyplot-User` on every `/v1` route; outside `ENVIRONMENT=development` it also requires the IAM-forwarded ID token (`AGENT_SERVICE_URLS`, `AGENT_ALLOWED_CALLERS`). To drive it from the plot page, run the API with `AGENT_ENABLED=true AGENT_SERVICE_URL=http://localhost:8001` (see `api/routers/agent.py`).

At the defaults only one run may start a minute, so a second "Create plot" within a minute waits in the run queue and the stream shows `status` events with `step: "queued"`. To iterate faster on your own machine, export `AGENT_RUNS_PER_MINUTE=60`. The theme toggle (`POST /v1/sessions/{sid}/versions/{version}/render {"theme": "dark"}`) never waits in the queue.

### Test

```bash
Expand Down
6 changes: 4 additions & 2 deletions agents/anyplot/agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,8 @@
* `root_agent` ("anyplot") is the only agent that talks to the user. Its policy is the
constant `static_instruction` (`policy.root_instruction`), never templated; the
`InstructionProvider` `session_context` adds only server-validated values (spec id,
library, reply language, dataset and binding status, plot versions), which ADK sends
library, reply language, dataset and binding status, plot versions and the latest
version's theme, so a change keeps it), which ADK sends
as a marked instruction block after the static prefix. Catalogue text such as the
spec title started as a public issue, so it never enters that block: the root reads
it fenced as `<spec_text>` through `get_spec_brief`. Its tools are the four session
Expand Down Expand Up @@ -86,7 +87,8 @@ async def session_context(context: ReadonlyContext) -> str:
lines.append(f"- Bindings: incomplete; missing roles: {missing}; {len(check.errors)} invalid")
versions = [version for version in services.versions.all(session_id) if version.library == view.library]
if versions:
lines.append(f"- Plot versions: {len(versions)}; latest result: {versions[-1].result.status}")
latest = versions[-1]
lines.append(f"- Plot versions: {len(versions)}; latest result: {latest.result.status}, theme {latest.theme}")
else:
lines.append("- Plot versions: none yet")
return "\n".join(lines)
Expand Down
25 changes: 15 additions & 10 deletions agents/anyplot/code/export.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,37 +5,42 @@
# Adapted by anyplot.ai from scatter-basic (matplotlib 3.11.2) for your data.csv;
# run: ANYPLOT_THEME=light python plot.py

The catalogue's four-line header and title rule do not apply to user plots
(docs/concepts/agent-network.md, "Fencing and exported code"). Exporting an export
replaces its header instead of stacking a second one, so `export_code` is idempotent.
The spec id, library and version are checked against strict patterns, because they
become source code.
The run line names the theme the run rendered (`PipelineArgs.theme`), so a user who
asked for a dark plot and follows the header gets the dark one; the code itself still
reads `ANYPLOT_THEME` and defaults to light. The catalogue's four-line header and
title rule do not apply to user plots (docs/concepts/agent-network.md, "Fencing and
exported code"). Exporting an export replaces its header instead of stacking a second
one, so `export_code` is idempotent. The spec id, library, version and theme are
checked against strict patterns, because they become source code.
"""

import re


HEADER_PREFIX = "# Adapted by anyplot.ai from "
RUN_LINE = "# run: ANYPLOT_THEME=light python plot.py"
RUN_LINE = "# run: ANYPLOT_THEME={theme} python plot.py"
EXPORT_THEMES = ("light", "dark")

_SPEC_ID = re.compile(r"[a-z0-9]+(?:-[a-z0-9]+)*")
_LIBRARY = re.compile(r"[a-z][a-z0-9]*")
_VERSION = re.compile(r"[0-9A-Za-z][0-9A-Za-z.+-]{0,31}")
_EXISTING_HEADER = re.compile(r"\A# Adapted by anyplot\.ai from [^\r\n]*\r?\n# run: [^\r\n]*\r?\n(?:\r?\n)?")


def attribution(spec_id: str, library: str, library_version: str | None) -> str:
def attribution(spec_id: str, library: str, library_version: str | None, theme: str = "light") -> str:
"""The two header lines plus the blank line after them."""
if not _SPEC_ID.fullmatch(spec_id):
raise ValueError(f"not a spec id: {spec_id!r}")
if not _LIBRARY.fullmatch(library):
raise ValueError(f"not a library id: {library!r}")
if library_version is not None and not _VERSION.fullmatch(library_version):
raise ValueError(f"not a library version: {library_version!r}")
if theme not in EXPORT_THEMES:
raise ValueError(f"not a theme: {theme!r}")
source = f"{library} {library_version}" if library_version else library
return f"{HEADER_PREFIX}{spec_id} ({source}) for your data.csv;\n{RUN_LINE}\n\n"
return f"{HEADER_PREFIX}{spec_id} ({source}) for your data.csv;\n{RUN_LINE.format(theme=theme)}\n\n"


def export_code(run_form: str, *, spec_id: str, library: str, library_version: str | None) -> str:
def export_code(run_form: str, *, spec_id: str, library: str, library_version: str | None, theme: str = "light") -> str:
"""The downloadable `plot.py`: the attribution header, then `run_form` unchanged."""
return attribution(spec_id, library, library_version) + _EXISTING_HEADER.sub("", run_form, count=1)
return attribution(spec_id, library, library_version, theme) + _EXISTING_HEADER.sub("", run_form, count=1)
Loading
Loading