Repository navigation
Fable 5.1 effort-marker correlation fails when a turn has more than one user message #237
Description
Activity
Second variant still live after #238 — different mechanism, same file
PR #238 fixed the
anchor_placementvariant (tool_result continuations pushing the anchor off the last user message). Reproduced again at 20:43:06Z and 20:44:04Z today, on a build that carries that fix:dist/index.js built 2026-09-18 20:32:22Z · contains anchor_placement errors at 2026-09-18 20:43:06Z, 20:44:04ZSo this is a genuinely separate path, not a regression or a stale bundle.
Mechanism
The two sides disagree about what a "boundary" is.
Markers are attached per HOST USER RECORD —
effort-history.ts:392:const token = transitionMarker(scope, info.id, effort) item.parts.push({ type: 'text', text: token })
Markers are validated per WIRE MESSAGE —
effort-history.ts:612:for (const message of consumed.messages) { if (message.transitions.length > 1) { throw new EffortMarkerCorrelationError( 'Multiple internal Fable 5.1 effort markers on one user boundary', ) }
consumeInternalMarkerswalksbody.messagesand groups every marker it strips by the wire message it came from. The Anthropic API requires strictly alternating user/assistant roles, so two consecutive host user records must collapse into one wire user message — at which point both of their transition markers land on a single boundary and the> 1guard fires.The invariant is broken by construction whenever two consecutive user records each carry an effort transition.
Why this session and not others
All four recorded occurrences are in one session:
ses_019303cc7ffe0QbSYzgqyZRDfj, a coordinator seat. That seat routinely receives consecutive user-role messages without an intervening assistant turn — peer session frames (<external-context source="sibling-session">) and subagent completion frames (<task state="completed">) arrive as ordinary user messages. Any two of those spanning an effort change produce the collapse.This is the multi-user-boundary hypothesis from the original report. It was wrong for
anchor_placement(that was tool continuations) and right for this one. Two variants, two mechanisms, both real — which is why fixing the first did not move this one.Occurrences
2026-09-15 09:29:50 2026-09-16 18:06:37 2026-09-18 20:43:06 ← post-#238 2026-09-18 20:44:04 ← post-#238Fix direction, not yet implemented
Two markers on one wire boundary is legitimate when they came from distinct host records that the lowering merged. The validator should reconcile against the request plan's ordered transition list rather than rejecting on count: N markers on one boundary are acceptable when they are a contiguous, correctly-ordered run of the plan, and the last one wins for the effective effort. A count check cannot distinguish that from genuine duplication, because it discards the identity the markers carry.
Fail-closed must be preserved for the cases it was built for: holes, reordering, mutation, scope mismatch, and a marker whose id is not in the plan.
Root cause, with the escalation dynamic that makes it cluster
Correcting my previous comment: the trigger is not merely "consecutive user records". A failed turn widens the next request's merge, so the fault is self-amplifying — which is why every occurrence sits in a run rather than alone.
The evidence chain
1. Markers attach per HOST record —
effort-history.ts:392const token = transitionMarker(scope, info.id, effort) item.parts.push({ type: 'text', text: token })
2. Markers are validated per WIRE message —
effort-history.ts:612for (const message of consumed.messages) { if (message.transitions.length > 1) { throw … }
consumeInternalMarkerswalksbody.messagesand groups by wire message.3. Consecutive host user records are routine on this seat — measured, not assumed. 26 consecutive-user pairs across the session, most at an identical timestamp:
2026-09-15 09:29:08 | user | low ← same instant 2026-09-15 09:29:08 | user | lowCoordinator seats receive peer s2s frames and subagent completion frames as ordinary user-role messages, so runs form without any user action.
4. A FAILED TURN LEAVES AN ASSISTANT RECORD WITH ZERO PARTS — measured:
msg_0b64313af0 parts=0 textparts=0 msg_0b6422fe90 parts=0 textparts=0 msg_0ab66355e0 parts=0 textparts=0An assistant record with no parts lowers to nothing, so it cannot separate the user records on either side of it.
5. Therefore the failure escalates. Observed at 09:29 on 09-15:
09:29:08 user low 09:29:08 user low 09:29:08 assistant ERR ← "correlation failed: expected N, found M" (0 parts) 09:29:50 user high 09:29:50 assistant ERR ← "Multiple internal effort markers" (0 parts)The first failure produces a contentless assistant. On the next request, that record contributes nothing to the wire, so the user records before and after it collapse together — now spanning a
low → highchange, so two distinct transition markers land on one wire boundary and the> 1guard fires. Every recorded occurrence is preceded by another failure or sits directly beside one (20:43:06 → 20:44:04, 58s apart).Why the guard is the wrong invariant
Two markers on one wire boundary is legitimate when they came from distinct host records the lowering merged. The markers carry
(scope, message-id, effort)— enough to tell a correctly-ordered run of the request plan from genuine duplication — but the count check discards that identity before it is consulted.Fix direction
Reconcile the stripped markers against the request plan's ordered transition list instead of rejecting on count:
- N markers on one boundary are valid when they form a contiguous, correctly-ordered subsequence of the plan, with the last one winning for effective effort.
- Keep fail-closed for what the guard was actually built for: holes, reordering, mutation, scope mismatch, and any marker whose id is absent from the plan.
A fix should be tested against the escalated shape (≥2 user records merged across an effort change), not just the single-boundary case — the first failure in a run is a different variant from the ones that follow it.
One inferred link, named
Steps 1, 2, 3, 4 and the clustering in 5 are measured from the tree and the host DB. The merge itself — that consecutive user-role records collapse into one wire message — is inferred from Anthropic's alternating-role requirement rather than observed directly, because the throw happens before the dump path and no request artifact survives. Pinning it needs a reproduction that lowers a known consecutive-user history and inspects
body.messages; that reproduction is also the natural regression test.PR #241 opened for the
Multiple internal … markers on one user boundaryvariant, implementing the fix direction above.Three changes: drop the per-message count guard (duplication and ordering are already enforced globally on the flat array), scope-check every transition on a boundary rather than
[0], and take the last transition for emission instead of the first.The third is what makes the first safe. Relaxing the guard alone would have converted this loud 400 into a silent wrong-effort request — the emission loop at
:724tooktransitions[0], so a merged boundary would have applied the stale effort. Mutation-proved: reverting to[0]yieldseffort: "high"where"max"is expected.Status of the two variants in this issue:
anchor_placement(tool-result continuations) — fixed in fix(opencode): accept effort-marked Fable 5.1 tool continuations #238, shippedMultiple internal … markers(merged user boundary) — fix(fable): accept effort markers on a merged user boundary #241correlation failed: expected N, found M— still open. It is the first failure in each observed run, and the one that manufactures the contentless assistant that widens the next merge. Worth treating as its own investigation rather than assuming fix(fable): accept effort markers on a merged user boundary #241 covers it; fix(fable): accept effort markers on a merged user boundary #241 stops the amplification but not the initiating failure.
Variant 3 root cause — the trimmed-prefix tolerance is unreachable in the two shapes that actually occur
correlation failed: expected N, found Mis the initiating failure in each run — the one that produces the contentless assistant that widens the next request's merge into the variant #241 fixes. All four fleet-wide occurrences:2026-09-12 08:47:23 ses_f7ed4e49… expected 1, found 0 2026-09-14 09:49:53 ses_019303cc… expected 4, found 1 2026-09-14 17:48:54 ses_f7ed4e49… expected 1, found 0 2026-09-15 09:29:08 ses_019303cc… expected 2, found 1Two sessions, two shapes, every one
found < expected— marker LOSS, never duplication.Marker loss is expected and already has designed tolerance.
effort-history.ts:658-688handles exactly this: the host trims a history prefix, the surviving markers must be an exact suffix of the plan, and consumed transitions fold into the retained baseline. That logic is correct. Neither occurring shape reaches it.Shape A —
expected 1, found 0(total trim)effort-history.ts:570-576:if (!hasCandidate) { if (!requestPlan) return { found: 0, inserted: 0 } if (requestPlan.markerCount !== 0) { throw new EffortMarkerCorrelationError( `Fable 5.1 effort marker correlation failed: expected ${requestPlan.markerCount}, found 0`, ) }
hasCandidateis false when no marker text appears anywhere in the body. This throw is unconditional onexpectedTransitionsand sits ~90 lines before the tolerance, so it short-circuits first.The result is an ordering defect: partial trimming is tolerated, total trimming is rejected. Trim 3 of 4 markers and the suffix logic accepts it; trim the 4th as well and the same legitimate compaction becomes a 400. Nothing about a full trim is less valid than a partial one — with
expectedTransitionspresent,trimmedPrefix === expectedTransitions.lengthis a well-defined case meaning "every transition was consumed, fold them all into the baseline".Shape B —
expected 4, found 1/expected 2, found 1(partial trim, no plan)These reach
:690, the exact path:} else { if (found !== requestPlan.markerCount) { throw … }
The
elsemeansexpectedTransitionswas null, i.e.resolveExpectedPlanreturned null becauseresolvedPlanwas undefined. So the request carried a valid plan header but the tracker could not resolve it back to a plan object — and without the plan there is nothing to compare a surviving suffix against, so any loss is fatal.OpenCodeEffortPlanTracker.resolveHeader(:783) resolves by re-encoding every tracked plan and string-comparing to the header. That only fails if the entry is gone from the map. Two candidate mechanisms, neither yet confirmed — this is where the investigation stands, not a conclusion:- Eviction.
MAX_TRACKED_EFFORT_PLANS = 1024, LRU by insertion order.markHeadersrefreshes recency, so a retried plan survives; a busy multi-project process is the scenario to check. clear()on the same message id.index.ts:5374-5387— whenmarkOpenCodeEffortTransitionsreturns no plan, the hook clears the entry for the current user message. If the hook runs a second time for a message whose plan was already recorded and the second pass produces no plan, the recorded entry is dropped while a header referencing it is already in flight.
Mechanism 2 is the one I would test first: the record and clear are the two arms of the same
if, and a coordinator seat re-running the messages hook across consecutive user records is precisely the traffic shape both affected sessions share.Why this is worth fixing beyond the 400 itself
Variant 3 manufactures the contentless assistant (
parts=0) that variant #241's merge amplification feeds on. #241 stops the amplification; this stops the ignition.Fix direction
- Shape A: move the
!hasCandidatecount check after the tolerance, or give it the sameexpectedTransitionsawareness — a full trim with a resolvable plan is thetrimmedPrefix === lengthcase and should fold all transitions into the baseline, not throw. Keep the throw for the genuinely untrusted case (markerCount > 0with no resolvable plan). - Shape B: confirm which mechanism drops the entry before changing anything. If it is
clear(), the fix is to not clear an entry that a header in flight may still reference; if it is eviction, the bound needs to account for concurrent in-flight requests.
Both need a regression test that lowers a trimmed history and asserts the request succeeds with the folded baseline, rather than asserting the error text.
- Eviction.
PR #242 opened for variant 3 — all three variants in this issue now have a fix in flight or shipped.
variant mechanism status anchor_placementtool-result continuation displaces the anchor from the last user message #238, shipped Multiple internal … markersconsecutive host user records merge into one wire boundary #241 correlation failed: expected N, found Mtrim tolerance unreachable — total trim short-circuits it; overwritten plan makes it unresolvable #242 The three are causally linked, not merely co-located. Variant 3 fires first and leaves a
parts=0contentless assistant; that record lowers to nothing, so it cannot separate the user records around it, and the next request's merge widens into variant 2. Fixing only the amplification would have left the ignition intact.Two things #242's investigation settled that are worth recording here:
A growing tool loop alone does not change the plan. The probe showed P1 == P2 across iterations. The prefix trim is what re-records the slot for the same
messageId— so the plan-overwrite window opens on compaction, not on tool depth. My earlier framing on this issue was too loose about that.Eviction and
clear()were ruled out with evidence, not assumed away. Eviction cannot be it because the in-flight plan is the newest andrecord+markHeadersboth refresh recency;clear()cannot be it because the transform returns a plan rather than null for a single Fable user.Leaving this issue open until #241 and #242 both land and a merged-boundary turn is observed completing in production.
@ualtinok, following up. Re-checked against current
main(4a01612). All three variants are still present, and each has a fix PR open:variant on mainnowfix tool-continuation anchor strict anchor-must-be-last-user-message check #238 multiple markers on one merged boundary packages/opencode/src/effort-history.ts:612if (message.transitions.length > 1), and emission usestransitions[0](:617,:722)#241 correlation failed: expected N, found M:570unconditionalif (!hasCandidate)throw on a full trim, and no plan history (MAX_TRACKED_EFFORT_PLAN_HISTORYabsent)#242 effort-history.tshasn't changed since these were filed; v1.23.0 didn't touch it.Production evidence for #241: it has been on my live host since 09-19. The shape that used to fail (two consecutive user records with an effort change between them) has since been seen completing with the fix in place.
Merge-order note: #238 and #242 both edit the same trailing test in
index.test.tsfor different reasons. #238 adds a misplaced-anchor case that endsexpect(messagesCalled).toBe(false). #242 makes the preceding full-trim case succeed, somessagesCalledis alreadytrueby then. Whichever merges second needs amessagesCalledreset between the two cases. I'll rebase that one once the first lands. #238's red CI is the upstream watcher crash fixed in #246, not #238 itself.I'll keep this issue open until all three fixes land and the merged-boundary case is seen completing in production.
Reopening because #238 and #241 fix tool-result anchor placement and merged user boundaries, but do not prove full marker-and-anchor loss is a valid prefix trim. The bounded plan-history part of #242 landed; its unanchored full-trim branch did not. A retained earlier user/assistant pair with all marker text removed is a concrete non-prefix counterexample, covered by the fail-closed regression in packages/opencode/src/tests/effort-history.test.ts. #242 was closed as superseded, not fully merged. Keeping this issue open until a surviving-boundary proof or production evidence supports a safe resolution.
Agreed: the counterexample holds. If a downstream transform strips all marker text but keeps the earlier user/assistant pair, #242's full-trim branch would fold every transition into the baseline and send the wrong effort. Nothing proves the messages were actually trimmed. Failing closed is the right call there.
Our live seat is moving to main in the next few hours, after its scoped-custody migration. From then on, every occurrence will log
missing_all_markers(retainedMessageCount,lastUserMessageIndex, block counts, plan metadata). I'll post those records here as they come in, so we can tell a genuine full trim apart from downstream stripping.
Symptom
A Fable 5.1 turn dies with a user-visible 400 carrying an internal message:
Recurring, not a one-off — 8 occurrences across 3 sessions in one week, in four shapes, all
EffortMarkerCorrelationError:Fable 5.1 effort marker correlation failed: expected N, found MMultiple internal Fable 5.1 effort markers on one user boundaryMissing or invalid internal Fable 5.1 effort anchorRoot cause
buildEffortRequestPlanpushes the anchor onto the user message that is current when the messages hook runs (effort-history.ts:409,currentUser.parts.push(...)), and stamps that message's id asanchor.boundary.Validation at lowering time then requires the anchor to sit on the last user message:
The invariant assumed is "exactly one user message at the current boundary." That holds for a plain chat turn and breaks as soon as the host puts a second user message at the same boundary — after which the anchor is on an earlier user message and the check fails.
Evidence
From the failing session (ids abbreviated):
Both failures are the second provider call within a turn (tool continuation), in a turn whose boundary carries two user messages. Neither injected message is flagged
syntheticorignored— they are ordinary user messages as far as the plugin can see.The
Multiple internal Fable 5.1 effort markers on one user boundaryvariant is the same root cause seen from the other side: more than one user message at a boundary means more than one marker-bearing candidate.Why it matters
The fail-closed itself is behaving as designed — refusing before cache shaping, signing and transport is the right call versus shipping a bad prefix. Two things make it worth fixing anyway:
effortMarkerFailureResponse(index.ts:322) synthesises a 400 for the user, and the throw happens before the dump path, so there is no log line and no dump artifact. Reconstructing the above required the host database. At minimum this path should log the mismatch it refused on (anchor boundary, last user id, marker count) — a fail-closed nobody can diagnose gets worked around instead of fixed.Not fixed by f9c867d
f9c867dbounds and orders desktop notice identities. These injected user messages are not desktop notices, and the strict check is still present on currentmain(effort-history.ts:639).Suggested direction
Tolerate user messages appended after the plan was built, as long as they carry no markers — i.e. require the anchor to be on the last user message the plan knew about, rather than the last user message in the lowered history. Holes, reordering, mutation, and scope/digest mismatch should keep failing exactly as they do now.
Environment: this is heaviest in multi-seat setups where session-to-session frames and subagent completion frames are delivered as ordinary user messages, but any host that appends a second user message at a turn boundary will hit it.