You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
B1, structured answers through output_config.format. On Claude, the adapter plan, the reviewer verdict and the scope and dataset judge's verdict now answer under a structured-output format instead of a forced tool call. The format carries the response schema made strict by the SDK's transform_schema, so constrained decoding enforces the JSON shape, types, enums and required fields. The existing repairs still mend the lengths and counts a grammar cannot enforce. Unlike a forced tool call, the format lets a thinking model think before it answers.
B2, tolerant edit matching, never fuzzy. The edit applier tries three tiers in order: exact, then trailing whitespace before line breaks ignored, then whole lines whose words agree once indentation and runs of spaces are ignored, with the replacement re-indented. The first tier with any match decides, and the unique-match and protected-region checks stay. A find that only matches the code before an earlier edit of the same plan fails with its own kind, stale_match, and its own repair line. The eval record counts the edits each tier applied in edit_tolerant. A theme token is now a value picked per theme from literals and earlier tokens only, so box-basic's data-driven _tight_text_color is plot code the adapter may change.
B3, the module allowlist agreed with the owner. Adapted code may now import the pure computing modules warnings, typing, calendar, decimal, fractions, zoneinfo, enum, dataclasses, operator, bisect, heapq, numbers and cmath, plus scipy.fft, scipy.linalg and scipy.constants. locale, random, os beyond getenv, sys, pathlib, string and all I/O stay banned. Each admitted module was walked for names that read, write, fetch or evaluate, and those are banned. The ESCAPE_ATTRIBUTES set and the private-name rule close re-export routes such as collections._sys.modules[...], statistics.sys and enum.bltns, and private helpers such as dataclasses._FuncBuilder. A banned import's repair line names the module and its alternative. The run record's banned_imports field labels each one from a closed vocabulary of module names, and anything else as ?, so no word the model writes as an import reaches the log.
B4, per-agent Claude thinking for measurements.AGENT_ROOT_THINKING_BUDGET and AGENT_ADAPTER_THINKING_BUDGET (0, off, by default; at most 4,096) turn on adaptive thinking for that agent and add the budget to its output cap. The eval harness sets both with --thinking-budget ROOT,ADAPTER, and the stamp and summary record it. With both at 0 no request changes.
This branch was stacked on the reviewer quality package and is rebased onto main after #12122 merged. It keeps main's cost-weighted budgets, JUDGE_WEIGHTS, the plot_shipped budget exemption and next_calls=2 unchanged.
Live smoke
On 2026-10-10 between 21:47 and 21:49 UTC, two eval cases ran on Vertex AI eu with the fake renderer from this branch, once plain and once with --thinking-budget 1024,2048. The runs used the branch's pre-rebase head 177e8df5a, before main's budget changes from #12122 were replayed underneath.
Vertex accepted output_config.format for claude-haiku-5-5. Every adapter, reviewer and root call finished with STOP, 4 plans parsed, and there were 0 answer-schema misses.
Thinking next to the format worked. The run used 6,979 thought tokens, the root's replay after the tool result was accepted, and the second review ran.
Arm
Cost for 2 runs, cold cache
p50 latency
Plain
$0.018
12.6 s
--thinking-budget 1024,2048
$0.0195
31.7 s
The latency figures come from a tiny sample.
A main-side fact, not a run of this branch: at 23:20 UTC on 2026-10-10, the plot_shipped budget exemption from #12122 was verified live on Claude through the local stack on main 5fd2c08. The tool-less closing call succeeded with tool_use and tool_result blocks in its history, and the reply arrived.
Known gaps (not in this PR)
Third-party re-exports such as matplotlib.Path, which is pathlib.Path, remain open.
statistics.random as a route to a random number generator remains open.
Private helpers of numpy, pandas, scipy, matplotlib and the other third-party packages were not walked, apart from numpy's _datasource. A module reached through an attribute chain, such as [np.lib][0]._other, can still reach any private name the denylist does not hold. Banning every private attribute on any object would close this, but it would refuse _legend and _axinfo in 2 catalogue specs, so it is an owner decision.
Plan
The quality levers of 2026-10-10 (session research), lever 4.
Test plan
uv run ruff check .
uv run ruff format --check .
uv run mypy api core agents
uv run pytest tests/unit/agents -q (4,174 passed)
uv run python -m tools.changelog check --base origin/main
Measurement arm on the next throwaway renderer: thinking off versus root 1,024 / adapter 2,048 against the committed baseline
…nfig.format
The adapter plan, the reviewer verdict and the judge verdict now ask Claude for
structured outputs instead of a forced tool call. VertexClaude adds
output_config.format next to ADK's effort when a request carries a response
schema and no tools; the judge sends the same format. The schema is made strict
by the SDK's own anthropic.transform_schema: additionalProperties false on every
object, and the keywords constrained decoding cannot enforce (minLength,
maxLength, maxItems, defaults) moved into the field descriptions as hints.
Branch A's repairs (repair_plan, repair_verdict) keep mending the lengths,
counts and the ok/defects rule after the strict check.
Why the format and not a strict forced tool: on Haiku 5.5 a forced tool call
skips thinking, so the adapter could never think before its plan; Opus 5.5,
Sonnet 5.5 and Fable 5.1 refuse a forced tool_choice outright, and accept the
format. The settings still allow only claude-haiku-5-5, now because the agents
run with thinking disabled. The judge treats a stop reason other than end_turn
(a refusal, a cut-off) as a failed attempt.
Not smoke-tested on Vertex AI (no model calls from this branch): one eu call
should confirm that Vertex forwards output_config.format for claude-haiku-5-5.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…s data-driven lines to the adapter
The rerun of spike X on main recorded 22 edit-apply failures by kind:
drift:theme_token 13, protected:placeholder 5, zero_match 4.
drift:theme_token: 12 of the 13 were box-basic matplotlib's full-file repair
rewriting `_tight_text_color = INK if THEME == "dark" else
IMPRINT_PALETTE[tightest_idx]`. That line follows the data, so the rule was too
strict, not the model. A theme token is now an assignment that tests THEME and
is built only from literals, THEME and earlier tokens; of the 2,745
theme-testing assignments in the 650 matplotlib and seaborn files, this is the
only one the narrower rule releases (a test pins that). The drift check also
compares protected statements as syntax trees, so a full file that re-states a
token with the same value in other quotes, spacing or comment passes, and the
drift line now says to copy the token's line unchanged.
protected:placeholder: the failure line names the df = load_user_data() line,
says to keep it out of find and replace, and names the line below it to anchor
an insertion on; adapter.md says the same and defines a theme token precisely.
zero_match: edits now match exact first, then with the blanks before line
breaks ignored (CRLF read as LF), then as whole lines whose words agree once
indentation and runs of spaces are ignored, with the replacement re-indented
by the one shift between find and code. The first tier with any match decides
and keeps the unique-match and protected-region checks; nothing is fuzzy. A
find that matches only the code before an earlier edit of the same plan is a
new kind, stale_match, with its own repair line. The pipeline_check line and
the eval record count the edits each tolerant tier applied (edit_tolerant).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…d its alternative
Owner decision 2026-10-10: the validator admits the pure computing modules of
the standard library (warnings, typing, calendar, decimal, fractions,
zoneinfo, enum, dataclasses, operator, bisect, heapq, numbers, cmath) and
scipy.fft, scipy.linalg and scipy.constants. locale, random, sys, pathlib,
string, io, os beyond getenv, sklearn.datasets, scipy.io, scipy.datasets and
statsmodels.formula stay banned.
Each admitted module was walked (Python 3.13.12, scipy 1.18.1) for names that
read, load, save, open, fetch or evaluate, and for the modules it re-exports;
the result is recorded next to COMMON_IMPORTS. What the walk found is banned:
warnings.formatwarning/showwarning/warn_explicit (linecache reads, file
writes), zoneinfo.reset_tzpath/available_timezones/ZoneInfo.from_file,
typing.get_type_hints/ForwardRef (they evaluate strings),
operator.attrgetter/methodcaller (getattr by string), enum.global_enum. The
walk also showed that an allowed module can hand out a module the import rules
keep out: enum.bltns is builtins, dataclasses re-exports inspect and has a
private _FuncBuilder that runs exec, and statistics.sys.modules or
collections._sys.modules reach every module, which passed the validator
before this change too. ESCAPE_ATTRIBUTES (sys, os, builtins, bltns, modules,
inspect, subprocess, importlib, pathlib, shutil, socket, linecache, CodeType,
FunctionType and the names above) are now banned as attributes on any object
and as from imports, and so is a private name of anything an import bound.
On the 650 normalised catalogue files no status and no rule set changed.
A banned-import message now names the module and what to use instead
(IMPORT_ALTERNATIVES, by longest known prefix; a reader or loader call says
the data comes only from df = load_user_data()). The finding carries the
module path (letters, digits, dots and underscores, at most 40 characters,
else ?), the pipeline_check attribution line lists them, and the eval record
counts them per module (banned_imports), so the next run shows what Haiku
reaches for: the rerun's 17 banned-import rejections always came with a
banned-call, which fits io plus pd.read_csv, but the log could not say.
adapter.md names the allowed set; the design doc's validator passage lists the
new modules and bans.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
AGENT_ROOT_THINKING_BUDGET and AGENT_ADAPTER_THINKING_BUDGET (0, off, by
default; at most 4,096) turn on Claude's thinking for the root or the adapter.
Claude Haiku 5.5 takes no thinking budget of its own (budget_tokens is a 400),
so make_content_config asks for ADK's adaptive mode (thinking_budget=-1, sent
as thinking adaptive with display summarized), the depth follows the kind's
effort, and the budget is added to the kind's output cap, so the answer keeps
its whole measured headroom however much of it the thinking uses.
allow_full_file adds the adapter's budget to the full-file cap again; 4,096 is
the largest value that keeps that cap (20,480) below the Anthropic SDK's limit
for a request that does not stream. The adapter can think because its answer
is now a structured-output format, not a forced tool call. The summary comes
back as thought parts, which the answer guard, ADK's output parsing and the
stream translator skip; thinking tokens are already counted and priced as
output. With both at 0 every request is byte-identical to before.
The settings refuse a budget on Gemini, which thinks by level. The harness
takes --thinking-budget ROOT,ADAPTER, writes both variables, records them in
the stamp as thinking_budget, and the Markdown summary names a measurement.
Risk for the root, recorded in its setting: it replays its thinking block
after a tool result, and an account that enforces the history-editing check of
preserved thinking may refuse that call.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…rdening
changelog.d/agents-adapter-hardening.md covers the structured-output answers,
the tolerant edit matching, the narrower theme-token rule and the AST drift
check, the placeholder repair line, the admitted computing modules and the
closed re-export routes, and the thinking measurement. agents/README.md and
the design doc's harness passage list the new record fields (edit_tolerant,
banned_imports) and the stamp's thinking budgets.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…laude cut-offs
The banned-import label on the pipeline_check line and in the eval record
kept any identifier the model wrote, so a word from the data written as an
import path reached the attribution log. validate.module_label now logs
the longest prefix of the path in LOGGED_MODULES (the standard library's
top-level names, the allowlists, the denied modules, the IMPORT_ALTERNATIVES
keys and a short list of third-party roots), and anything else as '?'.
import_alternative shares the prefix walk, and matrix.module_path re-checks
a logged label with the same function instead of a character pattern.
FakeAnthropic now cuts a structured answer stopped at max_tokens before its
closing quotes and brackets, so the Claude arms of the cut-off tests send
the truncated JSON text output_config.format leaves at the limit. The
adapter prompt no longer suggests enum and dataclasses, whose usual forms
need the class statement the validator bans. The spike A-prime row, the
service-flow comments and the module-label prose follow the format, and the
new catalogue sweep in test_regions skips without plots/ like its siblings.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…literal-aware tolerant matching
Applies the four findings of the first Copilot review on #12124.
The validator now checks every component of an import path below its root
against the escape names and for a leading underscore, so
`import scipy.fft._backend as backend` and `from scipy.fft._pocketfft import r2c`
are refused: an alias hid such components from the attribute check. The root
stays with the allowlist and keeps its message. calendar's LocaleTextCalendar,
LocaleHTMLCalendar and different_locale call `_locale.setlocale`, and
calendar.main reaches them through its `-L` option, so the four names join
ESCAPE_ATTRIBUTES. A sweep of the 650 matplotlib and seaborn catalogue files
finds no file the new rules touch.
The whitespace tier of the edit applier keeps a string literal's text verbatim
while it collapses the blanks between tokens, so `"North America"` no longer
matches `"North America"`, and it never matches a line that starts inside a
multi-line literal. Stale-match detection now looks for the find in the
pre-plan code with the same tiered locator, so a stale find written with other
blanks gets the stale-match repair line instead of the zero-match one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Copilot review, round 1: all four findings were real and are applied in c233ca6.
Calendar locale setters (applied).LocaleTextCalendar, LocaleHTMLCalendar and different_locale call _locale.setlocale, so they are now in ESCAPE_ATTRIBUTES. The walk also found calendar.main, which reaches the same classes through its -L option, so it is banned too.
Private and escape components in dotted imports (applied). Every component of an import path below its root is now checked for a leading underscore and against ESCAPE_ATTRIBUTES. import scipy.fft._backend as backend and from scipy.fft._pocketfft import r2c are refused. The root stays with the allowlist and keeps its message. None of the 650 matplotlib and seaborn catalogue files is affected.
Stale detection only in the exact tier (applied). The pre-plan code is now searched with the same tiered locator, so a stale find written with other blanks gets the stale_match repair line.
Whitespace tier collapsing blanks inside string literals (applied). The tier now collapses blanks only between tokens and keeps a literal's text verbatim. "North America" no longer matches "North America", and a line that starts inside a multi-line literal is never matched this way.
Each fix has tests, and the design doc and the changelog fragment describe the new rules.
Tolerant edits omitted from mixed-success failure attribution
agents/anyplot/pipeline.py:780
edit_tolerant is emitted only after the entire edit plan succeeds. If one edit is applied by a tolerant tier and a later sibling edit fails, AppliedPlan.tolerant contains that application but the edits_failed attribution branch omits it, so the new eval metric undercounts exactly the mixed-success plans it is intended to measure. Include the counter in the failure attribution too.
… by name, tolerant counts on failed plans
Applies the second Copilot review on #12124.
The private-name rule needs the import resolver, and an expression defeats
it: `c = [calendar][0]; c._locale.setlocale(...)` and
`d = [dataclasses][0]; d._FuncBuilder` passed. A name bound by `import X`
may now only be the receiver of an attribute (new rule `module-value`), which
none of the 650 matplotlib and seaborn catalogue files breaks. A module
reached through an attribute chain or a `from` import cannot be told from a
plain value without importing it, so the dangerous private helpers a walk of
the allowed modules found join ESCAPE_ATTRIBUTES and are banned by name on
any object: `_sys`, `_locale`, `_random`, `_tzpath`, `_common`,
`_getcategory`, `_FuncBuilder` and numpy's `_datasource`, whose `open` reads
files. The test that returns `subprocess` from a function now expects the
`module-value` finding next to the banned import.
The `edits_failed` attribution now carries `edit_tolerant`, and a drifted
plan keeps its tolerant tiers, so the eval record counts the tolerant matches
of a plan that another edit or a drift failed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Copilot review, round 2: both findings were real and are applied in 90a9519. This was the last review round for this PR.
Private-name guard bypassed through a laundered module (applied). A name bound by import X may now only be the receiver of an attribute, under the new rule module-value, so [calendar][0] and [dataclasses][0] are refused. None of the 650 matplotlib and seaborn catalogue files uses a module as a value. A module reached through an attribute chain or a from import cannot be told from a plain value without importing it. So the dangerous private helpers a walk of the allowed modules found are now banned by name on any object: _sys, _locale, _random, _tzpath, _common, _getcategory, _FuncBuilder and numpy's _datasource, whose open reads files. The corpus has the laundered forms. Third-party private helpers beyond _datasource were not walked, and the PR body lists that under known gaps.
Tolerant edits missing from the failure attribution (applied). The edits_failed line now carries edit_tolerant, and a drifted plan keeps its tolerant tiers. The eval record therefore counts tolerant matches in plans that another edit or a drift failed.
…ening
One conflict, in the regression-harness section of
docs/concepts/agent-network.md: main's fixture-case and harness bullets
(the 23 stress cases, the `stress` selector and `--static`) are kept, and
the harness bullet keeps this branch's `edit_tolerant` and `banned_imports`
record fields. matrix.py, pipeline.py, report.py and the README merged
cleanly with both sides intact: `--thinking-budget` next to `--static`, and
`banned_imports` next to main's adapter profile cap.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
`main` was in ESCAPE_ATTRIBUTES, which bans a name as an attribute on any
object, so a column read as `df.main` drew a needless repair round. By the
lead's decision it moves to MODULE_ESCAPES, which is checked against the
resolved attribute path (`calendar.main`, also under an alias), `from`
imports (`from calendar import main as x`) and import paths. The
`module-value` rule keeps the calendar module from reaching it through an
expression the resolver cannot follow. A corpus case now asserts that
`df.main` passes; the changelog fragment and the design doc state the
narrower scope.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Two conflicts, both in docs/concepts/agent-network.md, resolved on main's
text with this branch's change on top:
- The scope judge row keeps main's retry back-off and the `needs_data`
verdict; the Claude judge answers under a structured-output format
(`output_config.format`), not a forced tool call.
- The fixture-case bullet is main's (the scope set is now
`agents/evals/scope.evalset.json`); the harness bullet keeps the
`edit_tolerant` and `banned_imports` record fields.
The judge's answer schema comes from the `Verdict` literal, so main's
`needs_data` verdict reaches the constrained format unchanged. This branch's
judge test pinned the old three-value enum and now expects the four values.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Merged at 00:23 UTC; the deploy-agents trigger built and deployed it (Cloud Build 9cabcb7e, SUCCESS with the token smoke): revision anyplot-agents-b9cabcb7e serves 100 % at the service URL, so the structured answers, tolerant edits and the module allowlist are live in the private service. The API side stays dark until the owner sets the user-id secret and the API trigger's service URL.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
output_config.format. On Claude, the adapter plan, the reviewer verdict and the scope and dataset judge's verdict now answer under a structured-output format instead of a forced tool call. The format carries the response schema made strict by the SDK'stransform_schema, so constrained decoding enforces the JSON shape, types, enums and required fields. The existing repairs still mend the lengths and counts a grammar cannot enforce. Unlike a forced tool call, the format lets a thinking model think before it answers.findthat only matches the code before an earlier edit of the same plan fails with its own kind,stale_match, and its own repair line. The eval record counts the edits each tier applied inedit_tolerant. A theme token is now a value picked per theme from literals and earlier tokens only, so box-basic's data-driven_tight_text_coloris plot code the adapter may change.warnings,typing,calendar,decimal,fractions,zoneinfo,enum,dataclasses,operator,bisect,heapq,numbersandcmath, plusscipy.fft,scipy.linalgandscipy.constants.locale,random,osbeyondgetenv,sys,pathlib,stringand all I/O stay banned. Each admitted module was walked for names that read, write, fetch or evaluate, and those are banned. TheESCAPE_ATTRIBUTESset and the private-name rule close re-export routes such ascollections._sys.modules[...],statistics.sysandenum.bltns, and private helpers such asdataclasses._FuncBuilder. A banned import's repair line names the module and its alternative. The run record'sbanned_importsfield labels each one from a closed vocabulary of module names, and anything else as?, so no word the model writes as an import reaches the log.AGENT_ROOT_THINKING_BUDGETandAGENT_ADAPTER_THINKING_BUDGET(0, off, by default; at most 4,096) turn on adaptive thinking for that agent and add the budget to its output cap. The eval harness sets both with--thinking-budget ROOT,ADAPTER, and the stamp and summary record it. With both at 0 no request changes.This branch was stacked on the reviewer quality package and is rebased onto main after #12122 merged. It keeps main's cost-weighted budgets,
JUDGE_WEIGHTS, theplot_shippedbudget exemption andnext_calls=2unchanged.Live smoke
On 2026-10-10 between 21:47 and 21:49 UTC, two eval cases ran on Vertex AI
euwith the fake renderer from this branch, once plain and once with--thinking-budget 1024,2048. The runs used the branch's pre-rebase head 177e8df5a, before main's budget changes from #12122 were replayed underneath.output_config.formatforclaude-haiku-5-5. Every adapter, reviewer and root call finished withSTOP, 4 plans parsed, and there were 0 answer-schema misses.--thinking-budget 1024,2048The latency figures come from a tiny sample.
A main-side fact, not a run of this branch: at 23:20 UTC on 2026-10-10, the
plot_shippedbudget exemption from #12122 was verified live on Claude through the local stack on main 5fd2c08. The tool-less closing call succeeded withtool_useandtool_resultblocks in its history, and the reply arrived.Known gaps (not in this PR)
matplotlib.Path, which ispathlib.Path, remain open.statistics.randomas a route to a random number generator remains open._datasource. A module reached through an attribute chain, such as[np.lib][0]._other, can still reach any private name the denylist does not hold. Banning every private attribute on any object would close this, but it would refuse_legendand_axinfoin 2 catalogue specs, so it is an owner decision.Plan
The quality levers of 2026-10-10 (session research), lever 4.
Test plan
uv run ruff check .uv run ruff format --check .uv run mypy api core agentsuv run pytest tests/unit/agents -q(4,174 passed)uv run python -m tools.changelog check --base origin/main🤖 Generated with Claude Code
https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M