Skip to content

Fable 5.1 effort-marker correlation fails when a turn has more than one user message #237

Description

@iceteaSA

Symptom

A Fable 5.1 turn dies with a user-visible 400 carrying an internal message:

Missing or invalid internal Fable 5.1 effort anchor

Recurring, not a one-off — 8 occurrences across 3 sessions in one week, in four shapes, all EffortMarkerCorrelationError:

count message
4 Fable 5.1 effort marker correlation failed: expected N, found M
2 Multiple internal Fable 5.1 effort markers on one user boundary
2 Missing or invalid internal Fable 5.1 effort anchor

Root cause

buildEffortRequestPlan pushes the anchor onto the user message that is current when the messages hook runs (effort-history.ts:409, currentUser.parts.push(...)), and stamps that message's id as anchor.boundary.

Validation at lowering time then requires the anchor to sit on the last user message:

// effort-history.ts:636-646
if (
  requestPlan.markerCount > 0 &&
  (anchors.length !== 1 ||
    anchorMessageIndex !== lastUserMessageIndex ||   // <-- this
    !anchor ||
    anchor.scope !== requestPlan.scope)
) throw new EffortMarkerCorrelationError('Missing or invalid internal Fable 5.1 effort anchor')

The invariant assumed is "exactly one user message at the current boundary." That holds for a plain chat turn and breaks as soon as the host puts a second user message at the same boundary — after which the anchor is on an earlier user message and the check fails.

Evidence

From the failing session (ids abbreviated):

13:48:37  user       "Do a work laptop sync, ORW as well"      <- operator prompt
13:48:37  user       <external-context source="sibling-session" …>   <- second user message, same second
13:48:37  assistant  ok
13:49:10  assistant  ERR   <- tool continuation, no new user message
13:49:43  user       <task id="…" state="completed">           <- another injected user message
13:49:43  assistant  ok
13:49:58  assistant  ERR   <- tool continuation

Both failures are the second provider call within a turn (tool continuation), in a turn whose boundary carries two user messages. Neither injected message is flagged synthetic or ignored — they are ordinary user messages as far as the plugin can see.

The Multiple internal Fable 5.1 effort markers on one user boundary variant is the same root cause seen from the other side: more than one user message at a boundary means more than one marker-bearing candidate.

Why it matters

The fail-closed itself is behaving as designed — refusing before cache shaping, signing and transport is the right call versus shipping a bad prefix. Two things make it worth fixing anyway:

  1. The trigger is benign. A trailing user message carrying no marker does not corrupt the effort timeline the plan describes, so the refusal costs a turn for nothing.
  2. The failure is diagnostically dark. effortMarkerFailureResponse (index.ts:322) synthesises a 400 for the user, and the throw happens before the dump path, so there is no log line and no dump artifact. Reconstructing the above required the host database. At minimum this path should log the mismatch it refused on (anchor boundary, last user id, marker count) — a fail-closed nobody can diagnose gets worked around instead of fixed.

Not fixed by f9c867d

f9c867d bounds and orders desktop notice identities. These injected user messages are not desktop notices, and the strict check is still present on current main (effort-history.ts:639).

Suggested direction

Tolerate user messages appended after the plan was built, as long as they carry no markers — i.e. require the anchor to be on the last user message the plan knew about, rather than the last user message in the lowered history. Holes, reordering, mutation, and scope/digest mismatch should keep failing exactly as they do now.

Environment: this is heaviest in multi-seat setups where session-to-session frames and subagent completion frames are delivered as ordinary user messages, but any host that appends a second user message at a turn boundary will hit it.

Activity

  1. iceteaSA commented on Sep 18, 2026

    @iceteaSA
    ContributorAuthor

    Second variant still live after #238 — different mechanism, same file

    PR #238 fixed the anchor_placement variant (tool_result continuations pushing the anchor off the last user message). Reproduced again at 20:43:06Z and 20:44:04Z today, on a build that carries that fix:

    dist/index.js built 2026-09-18 20:32:22Z · contains anchor_placement
    errors at          2026-09-18 20:43:06Z, 20:44:04Z
    

    So this is a genuinely separate path, not a regression or a stale bundle.

    Mechanism

    The two sides disagree about what a "boundary" is.

    Markers are attached per HOST USER RECORD — effort-history.ts:392:

    const token = transitionMarker(scope, info.id, effort)
    item.parts.push({ type: 'text', text: token })

    Markers are validated per WIRE MESSAGE — effort-history.ts:612:

    for (const message of consumed.messages) {
      if (message.transitions.length > 1) {
        throw new EffortMarkerCorrelationError(
          'Multiple internal Fable 5.1 effort markers on one user boundary',
        )
      }

    consumeInternalMarkers walks body.messages and groups every marker it strips by the wire message it came from. The Anthropic API requires strictly alternating user/assistant roles, so two consecutive host user records must collapse into one wire user message — at which point both of their transition markers land on a single boundary and the > 1 guard fires.

    The invariant is broken by construction whenever two consecutive user records each carry an effort transition.

    Why this session and not others

    All four recorded occurrences are in one session: ses_019303cc7ffe0QbSYzgqyZRDfj, a coordinator seat. That seat routinely receives consecutive user-role messages without an intervening assistant turn — peer session frames (<external-context source="sibling-session">) and subagent completion frames (<task state="completed">) arrive as ordinary user messages. Any two of those spanning an effort change produce the collapse.

    This is the multi-user-boundary hypothesis from the original report. It was wrong for anchor_placement (that was tool continuations) and right for this one. Two variants, two mechanisms, both real — which is why fixing the first did not move this one.

    Occurrences

    2026-09-15 09:29:50
    2026-09-16 18:06:37
    2026-09-18 20:43:06   ← post-#238
    2026-09-18 20:44:04   ← post-#238
    

    Fix direction, not yet implemented

    Two markers on one wire boundary is legitimate when they came from distinct host records that the lowering merged. The validator should reconcile against the request plan's ordered transition list rather than rejecting on count: N markers on one boundary are acceptable when they are a contiguous, correctly-ordered run of the plan, and the last one wins for the effective effort. A count check cannot distinguish that from genuine duplication, because it discards the identity the markers carry.

    Fail-closed must be preserved for the cases it was built for: holes, reordering, mutation, scope mismatch, and a marker whose id is not in the plan.

  2. iceteaSA commented on Sep 18, 2026

    @iceteaSA
    ContributorAuthor

    Root cause, with the escalation dynamic that makes it cluster

    Correcting my previous comment: the trigger is not merely "consecutive user records". A failed turn widens the next request's merge, so the fault is self-amplifying — which is why every occurrence sits in a run rather than alone.

    The evidence chain

    1. Markers attach per HOST record — effort-history.ts:392

    const token = transitionMarker(scope, info.id, effort)
    item.parts.push({ type: 'text', text: token })

    2. Markers are validated per WIRE message — effort-history.ts:612

    for (const message of consumed.messages) {
      if (message.transitions.length > 1) { throw … }

    consumeInternalMarkers walks body.messages and groups by wire message.

    3. Consecutive host user records are routine on this seat — measured, not assumed. 26 consecutive-user pairs across the session, most at an identical timestamp:

    2026-09-15 09:29:08 | user | low     ← same instant
    2026-09-15 09:29:08 | user | low
    

    Coordinator seats receive peer s2s frames and subagent completion frames as ordinary user-role messages, so runs form without any user action.

    4. A FAILED TURN LEAVES AN ASSISTANT RECORD WITH ZERO PARTS — measured:

    msg_0b64313af0  parts=0  textparts=0
    msg_0b6422fe90  parts=0  textparts=0
    msg_0ab66355e0  parts=0  textparts=0
    

    An assistant record with no parts lowers to nothing, so it cannot separate the user records on either side of it.

    5. Therefore the failure escalates. Observed at 09:29 on 09-15:

    09:29:08  user low
    09:29:08  user low
    09:29:08  assistant ERR   ← "correlation failed: expected N, found M"   (0 parts)
    09:29:50  user high
    09:29:50  assistant ERR   ← "Multiple internal effort markers"          (0 parts)
    

    The first failure produces a contentless assistant. On the next request, that record contributes nothing to the wire, so the user records before and after it collapse together — now spanning a low → high change, so two distinct transition markers land on one wire boundary and the > 1 guard fires. Every recorded occurrence is preceded by another failure or sits directly beside one (20:43:06 → 20:44:04, 58s apart).

    Why the guard is the wrong invariant

    Two markers on one wire boundary is legitimate when they came from distinct host records the lowering merged. The markers carry (scope, message-id, effort) — enough to tell a correctly-ordered run of the request plan from genuine duplication — but the count check discards that identity before it is consulted.

    Fix direction

    Reconcile the stripped markers against the request plan's ordered transition list instead of rejecting on count:

    • N markers on one boundary are valid when they form a contiguous, correctly-ordered subsequence of the plan, with the last one winning for effective effort.
    • Keep fail-closed for what the guard was actually built for: holes, reordering, mutation, scope mismatch, and any marker whose id is absent from the plan.

    A fix should be tested against the escalated shape (≥2 user records merged across an effort change), not just the single-boundary case — the first failure in a run is a different variant from the ones that follow it.

    One inferred link, named

    Steps 1, 2, 3, 4 and the clustering in 5 are measured from the tree and the host DB. The merge itself — that consecutive user-role records collapse into one wire message — is inferred from Anthropic's alternating-role requirement rather than observed directly, because the throw happens before the dump path and no request artifact survives. Pinning it needs a reproduction that lowers a known consecutive-user history and inspects body.messages; that reproduction is also the natural regression test.

  3. iceteaSA commented on Sep 18, 2026

    @iceteaSA
    ContributorAuthor

    PR #241 opened for the Multiple internal … markers on one user boundary variant, implementing the fix direction above.

    Three changes: drop the per-message count guard (duplication and ordering are already enforced globally on the flat array), scope-check every transition on a boundary rather than [0], and take the last transition for emission instead of the first.

    The third is what makes the first safe. Relaxing the guard alone would have converted this loud 400 into a silent wrong-effort request — the emission loop at :724 took transitions[0], so a merged boundary would have applied the stale effort. Mutation-proved: reverting to [0] yields effort: "high" where "max" is expected.

    Status of the two variants in this issue:

  4. iceteaSA commented on Sep 18, 2026

    @iceteaSA
    ContributorAuthor

    Variant 3 root cause — the trimmed-prefix tolerance is unreachable in the two shapes that actually occur

    correlation failed: expected N, found M is the initiating failure in each run — the one that produces the contentless assistant that widens the next request's merge into the variant #241 fixes. All four fleet-wide occurrences:

    2026-09-12 08:47:23  ses_f7ed4e49…  expected 1, found 0
    2026-09-14 09:49:53  ses_019303cc…  expected 4, found 1
    2026-09-14 17:48:54  ses_f7ed4e49…  expected 1, found 0
    2026-09-15 09:29:08  ses_019303cc…  expected 2, found 1
    

    Two sessions, two shapes, every one found < expected — marker LOSS, never duplication.

    Marker loss is expected and already has designed tolerance. effort-history.ts:658-688 handles exactly this: the host trims a history prefix, the surviving markers must be an exact suffix of the plan, and consumed transitions fold into the retained baseline. That logic is correct. Neither occurring shape reaches it.

    Shape A — expected 1, found 0 (total trim)

    effort-history.ts:570-576:

    if (!hasCandidate) {
      if (!requestPlan) return { found: 0, inserted: 0 }
      if (requestPlan.markerCount !== 0) {
        throw new EffortMarkerCorrelationError(
          `Fable 5.1 effort marker correlation failed: expected ${requestPlan.markerCount}, found 0`,
        )
      }

    hasCandidate is false when no marker text appears anywhere in the body. This throw is unconditional on expectedTransitions and sits ~90 lines before the tolerance, so it short-circuits first.

    The result is an ordering defect: partial trimming is tolerated, total trimming is rejected. Trim 3 of 4 markers and the suffix logic accepts it; trim the 4th as well and the same legitimate compaction becomes a 400. Nothing about a full trim is less valid than a partial one — with expectedTransitions present, trimmedPrefix === expectedTransitions.length is a well-defined case meaning "every transition was consumed, fold them all into the baseline".

    Shape B — expected 4, found 1 / expected 2, found 1 (partial trim, no plan)

    These reach :690, the exact path:

    } else {
      if (found !== requestPlan.markerCount) { throw … }

    The else means expectedTransitions was null, i.e. resolveExpectedPlan returned null because resolvedPlan was undefined. So the request carried a valid plan header but the tracker could not resolve it back to a plan object — and without the plan there is nothing to compare a surviving suffix against, so any loss is fatal.

    OpenCodeEffortPlanTracker.resolveHeader (:783) resolves by re-encoding every tracked plan and string-comparing to the header. That only fails if the entry is gone from the map. Two candidate mechanisms, neither yet confirmed — this is where the investigation stands, not a conclusion:

    1. Eviction. MAX_TRACKED_EFFORT_PLANS = 1024, LRU by insertion order. markHeaders refreshes recency, so a retried plan survives; a busy multi-project process is the scenario to check.
    2. clear() on the same message id. index.ts:5374-5387 — when markOpenCodeEffortTransitions returns no plan, the hook clears the entry for the current user message. If the hook runs a second time for a message whose plan was already recorded and the second pass produces no plan, the recorded entry is dropped while a header referencing it is already in flight.

    Mechanism 2 is the one I would test first: the record and clear are the two arms of the same if, and a coordinator seat re-running the messages hook across consecutive user records is precisely the traffic shape both affected sessions share.

    Why this is worth fixing beyond the 400 itself

    Variant 3 manufactures the contentless assistant (parts=0) that variant #241's merge amplification feeds on. #241 stops the amplification; this stops the ignition.

    Fix direction

    • Shape A: move the !hasCandidate count check after the tolerance, or give it the same expectedTransitions awareness — a full trim with a resolvable plan is the trimmedPrefix === length case and should fold all transitions into the baseline, not throw. Keep the throw for the genuinely untrusted case (markerCount > 0 with no resolvable plan).
    • Shape B: confirm which mechanism drops the entry before changing anything. If it is clear(), the fix is to not clear an entry that a header in flight may still reference; if it is eviction, the bound needs to account for concurrent in-flight requests.

    Both need a regression test that lowers a trimmed history and asserts the request succeeds with the folded baseline, rather than asserting the error text.

  5. iceteaSA commented on Sep 18, 2026

    @iceteaSA
    ContributorAuthor

    PR #242 opened for variant 3 — all three variants in this issue now have a fix in flight or shipped.

    variant mechanism status
    anchor_placement tool-result continuation displaces the anchor from the last user message #238, shipped
    Multiple internal … markers consecutive host user records merge into one wire boundary #241
    correlation failed: expected N, found M trim tolerance unreachable — total trim short-circuits it; overwritten plan makes it unresolvable #242

    The three are causally linked, not merely co-located. Variant 3 fires first and leaves a parts=0 contentless assistant; that record lowers to nothing, so it cannot separate the user records around it, and the next request's merge widens into variant 2. Fixing only the amplification would have left the ignition intact.

    Two things #242's investigation settled that are worth recording here:

    A growing tool loop alone does not change the plan. The probe showed P1 == P2 across iterations. The prefix trim is what re-records the slot for the same messageId — so the plan-overwrite window opens on compaction, not on tool depth. My earlier framing on this issue was too loose about that.

    Eviction and clear() were ruled out with evidence, not assumed away. Eviction cannot be it because the in-flight plan is the newest and record + markHeaders both refresh recency; clear() cannot be it because the transform returns a plan rather than null for a single Fable user.

    Leaving this issue open until #241 and #242 both land and a merged-boundary turn is observed completing in production.

  6. iceteaSA commented on Sep 23, 2026

    @iceteaSA
    ContributorAuthor

    @ualtinok, following up. Re-checked against current main (4a01612). All three variants are still present, and each has a fix PR open:

    variant on main now fix
    tool-continuation anchor strict anchor-must-be-last-user-message check #238
    multiple markers on one merged boundary packages/opencode/src/effort-history.ts:612 if (message.transitions.length > 1), and emission uses transitions[0] (:617, :722) #241
    correlation failed: expected N, found M :570 unconditional if (!hasCandidate) throw on a full trim, and no plan history (MAX_TRACKED_EFFORT_PLAN_HISTORY absent) #242

    effort-history.ts hasn't changed since these were filed; v1.23.0 didn't touch it.

    Production evidence for #241: it has been on my live host since 09-19. The shape that used to fail (two consecutive user records with an effort change between them) has since been seen completing with the fix in place.

    Merge-order note: #238 and #242 both edit the same trailing test in index.test.ts for different reasons. #238 adds a misplaced-anchor case that ends expect(messagesCalled).toBe(false). #242 makes the preceding full-trim case succeed, so messagesCalled is already true by then. Whichever merges second needs a messagesCalled reset between the two cases. I'll rebase that one once the first lands. #238's red CI is the upstream watcher crash fixed in #246, not #238 itself.

    I'll keep this issue open until all three fixes land and the merged-boundary case is seen completing in production.

  7. ualtinok commented on Sep 24, 2026

    @ualtinok
    Contributor

    Reopening because #238 and #241 fix tool-result anchor placement and merged user boundaries, but do not prove full marker-and-anchor loss is a valid prefix trim. The bounded plan-history part of #242 landed; its unanchored full-trim branch did not. A retained earlier user/assistant pair with all marker text removed is a concrete non-prefix counterexample, covered by the fail-closed regression in packages/opencode/src/tests/effort-history.test.ts. #242 was closed as superseded, not fully merged. Keeping this issue open until a surviving-boundary proof or production evidence supports a safe resolution.

  8. iceteaSA commented on Sep 24, 2026

    @iceteaSA
    ContributorAuthor

    Agreed: the counterexample holds. If a downstream transform strips all marker text but keeps the earlier user/assistant pair, #242's full-trim branch would fold every transition into the baseline and send the wrong effort. Nothing proves the messages were actually trimmed. Failing closed is the right call there.

    Our live seat is moving to main in the next few hours, after its scoped-custody migration. From then on, every occurrence will log missing_all_markers (retainedMessageCount, lastUserMessageIndex, block counts, plan metadata). I'll post those records here as they come in, so we can tell a genuine full trim apart from downstream stripping.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions