Skip to content

feat(app): agent chat page, result card and .adapt() button - #12114

Merged
MarkusNeusinger merged 6 commits into
mainfrom
feat/agent-chat-ui
Oct 10, 2026
Merged

MarkusNeusinger merged 6 commits into
mainfrom
feat/agent-chat-ui

Conversation

@MarkusNeusinger

Copy link
Copy Markdown
Owner

Summary

  • Adds the admin-only "Use with my data" chat page at /debug/agent?spec=&library=&language= (app/src/pages/AgentChatPage.tsx): data panel with a 200 KB counter, 20-row preview, a column dropdown for every spec role (series families get one dropdown per member), the chat thread, a progress timeline that follows the anyplot/1 stream, a queue-position counter in the chat's corner, .stop() through the cancel route, and library pills. Built only with VITE_ENABLE_AGENT_CHAT=true (off in app/cloudbuild.yaml, so production ships without the chunk) and opens only in local development or for a browser carrying the admin hint.
  • Adds the result card (app/src/sections/agent-chat/ResultCard.tsx) with the plot page's overlay actions (copy image, download PNG, open full size), a light/dark switch through the theme toggle route, the adapted plot.py with copy and download, data.csv download, the change list and residual notes, a refine composer and the quick-feedback slot; plus the .adapt() overlay button on the plot page, gated by build flag, admin hint and the BFF's eligibility route so public visitors never call the debug API.
  • Two small contract changes on the service side so the page can address versions and roles reliably: the plot event carries version (agents translator and BFF allowlist), and the dataset answer lists the spec's roles with a 422 that keeps the binding check's lines. Analytics events (enum properties only) are documented in docs/reference/plausible.md; a loopback mock of the BFF (app/scripts/agent-bff-mock.mjs) drives the whole flow without the API, the agents service or a model.

Plan

Design: docs/concepts/agent-network.md (sections "Frontend", "Result card", "Analytics"); this is the frontend PR of the phase-1 roadmap, built on the merged runtime core (#12111) and run queue (#12112).

Test plan

  • cd app && yarn lint && yarn fm:check && yarn type-check && yarn test (81 files, 750 tests) and yarn build with the flag on and off; the default build contains no agent client
  • uv run --extra dev ruff check . && ruff format --check ., mypy api core agents, pytest tests/unit tests/integration (7816 passed), changelog check, uv lock --check
  • Driven in the browser against the mock at 1280 px and 390 px, light and dark: parse, bindings incl. series families and a refused set, queued counter, progress timeline, result card actions (clipboard PNG, downloads named plot.py/data.csv, open full size, copy code), theme switch, repair round, refusal, error, stop then a new turn without a 409, ineligible pair; no horizontal scroll, 16 px gutter at 390 px
  • After merge: deploy-app and deploy-api Cloud Builds succeed; production stays unchanged (_VITE_ENABLE_AGENT_CHAT false, AGENT_ENABLED false)
  • Once the agents service is deployed: open /debug/agent?spec=scatter-basic&library=matplotlib as admin and run one "Create plot" end to end

MarkusNeusinger and others added 2 commits October 10, 2026 01:30
…() button

Admin-only chat page at /debug/agent?spec=&library=&language= over the
/debug/agent BFF: data panel (200 KB counter, parse, 20-row preview,
warnings, one dropdown per spec role with the server's defaults), chat
thread with the fixed refusal and errors with code and ref, a progress
timeline and queue counter from the anyplot/1 status events, Stop
(AbortController plus the cancel route) and the library pills.

ResultCard per version: the rendered PNG from its blob URL, a light/dark
switch through the theme toggle route, copy image (Clipboard API with a
download fallback), download PNG, open full size, plot.py in
CodeHighlighter with copy code, downloads of plot.py and data.csv, the
change list, residual notes, a refine composer and a feedback slot.

lib/sse.ts parses SSE over fetchWithAuth's body reader (pings, CRLF,
split chunks); hooks/useAgentSession.ts holds the session state machine.
The .adapt() overlay button shows only with VITE_ENABLE_AGENT_CHAT, the
admin hint DebugPage now sets after /debug/status, and an eligible pair.
Without the flag the chunk is not built. Analytics are enum-only
(agent_open, agent_data_parsed, agent_plot_rendered,
agent_guardrail_block, copy_code page agent_chat) and documented in
plausible.md. app/scripts/agent-bff-mock.mjs serves the BFF routes and a
scripted stream locally for browser verification.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…actions

Review fixes for the "Use with my data" chat page:

- The plot event carries `version`, the number the agents service stored
  the result under (PlotResult.version, PLOT_FIELDS, the BFF allowlist);
  the page addresses the artifact and theme toggle routes by it and counts
  plot events only against a server without the field.
- The dataset answer lists the spec's roles {name, kinds, required,
  variadic, description}; the data panel gives every role a dropdown and a
  variadic family one per member plus one for the next, kept contiguous.
  A refused binding set keeps the check's lines (at most 20) through the BFF.
- The page starts no turn, parse, binding change, library switch or theme
  render while another is in flight; Stop keeps the stream until the
  server's done (at most 30 s), so a late plot still arrives and the next
  turn does not meet 409 run_active.
- The session is deleted on pagehide with keepalive; a request without an
  answer is `unreachable`, reloads once for Cloudflare Access (the guard is
  shared with DebugPage) and then offers a reload link.
- An ineligible pair offers the other libraries as links.
- The mock BFF listens on loopback only, serves line-multi with a series
  family, the roles, the error lines and the version; README and design doc
  name AGENT_USER_ID_KEY and ADMIN_TOKEN for a local run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Copilot AI balanced review requested due to automatic review settings October 10, 2026 00:21
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The mock does not preserve version artifacts, language inputs can become inconsistent, and the SSE line guard is bypassable.

4 open findings
What changed in this PR

Adds the feature-flagged, admin-only agent chat UI, result cards, .adapt() entry point, and supporting agent/BFF contract updates.

Changes:

  • Adds the complete agent-chat frontend and local mock.
  • Adds server-provided plot versions, roles, and binding diagnostics.
  • Documents, tests, and deploy-gates the feature.
File Description
agents/​README.md Documents local chat setup and mock usage.
agents/​anyplot/​pipeline.py Assigns stored version numbers to plot results.
agents/​anyplot/​schemas.py Adds validated plot-result versions.
agents/​main.py Returns role metadata with parsed datasets.
agents/​stream.py Includes versions in plot events.
api/​routers/​agent.py Relays versions and sanitized binding errors.
app/​ARCHITECTURE.md Documents agent-chat frontend structure.
app/​Dockerfile Adds the chat build argument.
app/​cloudbuild.yaml Keeps the production chat build disabled.
app/​scripts/​agent-bff-mock.mjs Implements the local BFF mock.
app/​src/​global-config.ts Defines the build-time feature flag.
app/​src/​hooks/​useAgentEligibility.test.ts Tests eligibility gating.
app/​src/​hooks/​useAgentEligibility.ts Gates .adapt() eligibility requests.
app/​src/​hooks/​useAgentSession.test.ts Tests chat session behavior.
app/​src/​hooks/​useAgentSession.ts Implements the chat state machine.
app/​src/​lib/​agent.test.ts Tests the BFF client and protocol parsing.
app/​src/​lib/​agent.ts Adds the typed agent-chat client.
app/​src/​lib/​sse.test.ts Tests incremental SSE parsing.
app/​src/​lib/​sse.ts Adds fetch-based SSE parsing.
app/​src/​pages/​AgentChatPage.test.tsx Tests page gating and startup.
app/​src/​pages/​AgentChatPage.tsx Adds the agent-chat page.
app/​src/​pages/​DebugPage.test.tsx Tests admin-hint updates.
app/​src/​pages/​DebugPage.tsx Shares admin authentication state.
app/​src/​pages/​SpecPage.tsx Connects eligible plots to the chat.
app/​src/​routes/​index.tsx Registers the feature-gated route.
app/​src/​routes/​paths.test.ts Tests agent-chat URL generation.
app/​src/​routes/​paths.ts Adds the agent-chat path helper.
app/​src/​sections/​agent-chat/​ChatThread.tsx Renders conversation items and results.
app/​src/​sections/​agent-chat/​DataPanel.test.tsx Tests role-binding controls.
app/​src/​sections/​agent-chat/​DataPanel.tsx Adds data parsing, preview, and bindings.
app/​src/​sections/​agent-chat/​ProgressTimeline.tsx Displays queue and pipeline progress.
app/​src/​sections/​agent-chat/​ReloadHint.tsx Adds recoverable-error reload actions.
app/​src/​sections/​agent-chat/​ResultCard.test.tsx Tests result actions and refinement.
app/​src/​sections/​agent-chat/​ResultCard.tsx Adds plot, artifact, theme, and code controls.
app/​src/​sections/​agent-chat/​bindings.ts Builds binding rows for role families.
app/​src/​sections/​agent-chat/​files.ts Adds clipboard and download helpers.
app/​src/​sections/​agent-chat/​messages.ts Maps protocol errors to UI copy.
app/​src/​sections/​agent-chat/​styles.ts Defines shared chat styles.
app/​src/​sections/​spec-detail/​SpecDetailView.test.tsx Tests .adapt() visibility and interaction.
app/​src/​sections/​spec-detail/​SpecDetailView.tsx Adds the .adapt() overlay action.
app/​src/​utils/​adminAuth.ts Centralizes admin token and hint storage.
app/​src/​vite-env.d.ts Types the feature flag.
changelog.d/​agent-chat-ui.md Records the feature and contract changes.
docs/​concepts/​agent-network.md Updates the implementation status and design.
docs/​reference/​api.md Documents roles, versions, and binding errors.
docs/​reference/​plausible.md Documents agent-chat analytics.
docs/​reference/​repository.md Adds the new frontend structure.
tests/​unit/​agents/​runtime/​test_service_flow.py Tests roles and version sequencing.
tests/​unit/​agents/​runtime/​test_stream.py Tests role serialization and plot versions.
tests/​unit/​agents/​test_schemas.py Tests version validation rules.
tests/​unit/​api/​test_agent_router.py Tests BFF filtering and error relaying.

🧠 Review effort: Balanced


💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread app/scripts/agent-bff-mock.mjs
Comment thread app/src/lib/sse.ts
Comment thread app/src/pages/AgentChatPage.tsx Outdated
Comment thread app/src/sections/agent-chat/styles.ts Outdated
MarkusNeusinger and others added 3 commits October 10, 2026 02:33
…lain action CTA, mock versions snapshot their data

Copilot's four findings on #12114:

- The SSE parser checked the line limit only on the unfinished remainder of
  the buffer, so an oversized line whose newline arrived in the same chunk
  was parsed. The limit now applies to every line before it is parsed; two
  tests cover the terminated oversized line and a line at the limit.
- The chat page took `language` from the query independently of `library`,
  so a mismatched pair picked the wrong highlighter and a back link to a
  page that does not exist. The language now follows the library through
  `LIB_TO_LANG`, as the catalogue pairs them.
- `.create_plot()` used the filled hero CTA, which the style guide reserves
  for the landing hero. It is now the regular method-call action with a
  rule border; `ctaButtonSx` is gone.
- The mock's versions read the session's current dataset and bindings, so an
  earlier result card served the new data after a reparse. A version now
  keeps the dataset and bindings it was made from, and its code, CSV and
  later theme renders read from that snapshot.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
# Conflicts:
#	docs/concepts/agent-network.md
# Conflicts:
#	agents/README.md
#	docs/concepts/agent-network.md
#	tests/unit/agents/runtime/test_stream.py
@MarkusNeusinger
MarkusNeusinger merged commit 6eead4c into main Oct 10, 2026
13 checks passed
@MarkusNeusinger
MarkusNeusinger deleted the feat/agent-chat-ui branch October 10, 2026 07:34
MarkusNeusinger added a commit that referenced this pull request Oct 10, 2026
## Summary
- The tool-safety allowlist for the `plot_pipeline` result did not
include `version` (added to `PlotResult` in #12114 for the chat UI), so
every shipped plot reached the root as `invalid_result` and the root's
closing reply told the user the run had failed while the plot was
already on screen. Seen in the first local run through the
`/debug/agent` gate.
- The allowlist is now derived from `PlotResult.model_fields` plus the
tool-error shape, so a schema field can never silently break the root
again; a plugin test passes a shipped result with `version=1` through
the callback and asserts the schema's fields stay allowed.

## Plan
N/A (bug found while bringing the stack up locally; the eval harness
never checks the root's text).

## Test plan
- [x] `ruff check`, `ruff format --check`, `mypy agents`, `pytest
tests/unit/agents/runtime/test_plugins.py
tests/unit/agents/runtime/test_service_flow.py` (85 passed), changelog
check
- [ ] CI green; after merge the local run's root reply reads the shipped
result ("your plot is ready" instead of "internal error")

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants