Repository navigation
Add Session Trajectory Mini App: per-turn messages, tool calls, Token usage and timing - #15
Conversation
MyPrototypeWhat
left a comment
There was a problem hiding this comment.
感谢贡献!会话代链的拼接和"怎么判定当前会话"的说明都做得很用心,npm run check 和 CI 也都通过了。我用合成的会话目录和 runtime-state.sqlite 在本地把服务和页面跑了一遍(headless Chrome),下面几处是实际复现到的,具体位置见行内评论。
需要修改
server.mjs:985:判断用的是turn.records.length === 0,所以只要同一轮前面已经有用户消息,content: []的助手消息(比如stopReason: 'error'的失败回复)就不会生成任何行。复现时「消息 9」但「已显示 8 / 共 8 条记录」,那一轮的 Token 也找不到对应的行。- PR 描述里说孤儿快照"reported as an orphan rather than shown",但它还是会出现在「上下文代」下拉里(例如「第 5 代(400 条)」),选中后是空页;
messagesAll也把它算了进去,于是显示成「全程 409 · 跨 4 代」,实际只拼了 3 代、9 条。 - 截断提示写的是「仅显示前 8 MB」,代码和 README 都是 64 MB。另外
truncated的条件是head.read < head.size,活跃文件比 catalog 记录的长度长时(也就是正在追加写入的时候)也会变成 true。复现时 688 字节的文件也显示了「文件过大」。 - README 说
/api/runtime"no absolute path appears in the response",但它原样返回了argv(argv[0]就是 node 可执行文件的绝对路径)、execArgv、pid、ppid。建议去掉或者只留 basename,或者至少在 README 里如实说明。 openRuntimeStateDb()每次请求都新开一个DatabaseSync,用完没有close(),dispose()也不会关。我在本地数了一下,轮询过程中一直有 20–30 个 runtime-state.sqlite 句柄开着,dispose 之后也还在。AGENTS.md 要求 dispose 关闭所有打开的句柄,加个try/finally db.close()就行。
建议(可选)
- 缺少
history-catalog.json的会话会出现在会话列表里,但一打开就是 500(artifact_path_rejected)。不确定真实客户端里有没有这种会话,如果有,可以让活跃文件在没有 catalog 时也能读取。 history_relative_dir直接join到 sessions 根目录,没有做resolveArtifactPath那样的包含检查,也没有比对目录名解码出来的 id;README 里"与 manifest、数据库交叉校验"的说法和代码对不上。数据来自 Host 数据库,风险不大,补一个检查或改一下措辞就好。- README 说
MMC_TRAJECTORY_ROOT"overrides the whole path",但代码把它当作数据根目录,后面还会拼v2/sessions和v2/sqlite,措辞可以改一下。
| if (turn.records.length === 0) { | ||
| push({ kind: 'assistant', title: 'assistant', summary: '', text: null, tokens: carrierUsage() }); | ||
| } |
There was a problem hiding this comment.
这里判断的是整轮有没有记录,不是这条消息有没有记录。同一轮前面已经有用户消息时,content: [] 的助手消息(比如 stopReason: 'error')就一行都不会出,它的 usage 仍然会算进本轮 Token,但页面上找不到对应的行。改成 if (segmentOrder.length === 0) 应该就是原本的意思。
There was a problem hiding this comment.
你说得对,是我读错了层级。records.length 是整轮累积的,所以只要这一轮前面已经有记录,后面那条 content: [] 的助手消息就整条不落到台账里,而它的 usage 照样进了 Token 总量——页面上找不到对应行。
改成看这条消息自己的 segmentOrder:server.mjs:1073。
| const messagesAll = generations.length > 1 | ||
| ? generations.reduce((sum, entry) => sum + (Number.isInteger(entry.messageCount) ? entry.messageCount : 0), 0) | ||
| : rows.length; |
There was a problem hiding this comment.
generations 里包含孤儿快照(没有被代链拼进来的),这里把它的 messageCount 也加进了「全程」。页面上的 generationOptions() 也会把它列成可选的一代,选中后是空的,和 PR 描述的"reported as an orphan rather than shown"不一致。可以在传给页面之前,按 lineage.chain 过滤一下。
There was a problem hiding this comment.
确认存在,按 lineage.chain 过滤了。
过滤放在 buildTrajectoryPayload 里面,按 orphans 的 fileName 匹配而不是在路由里先筛一遍再传进去(server.mjs:1283-1286)——这样调用方不管传进来的是 catalog 还是 chain 都不会再错,也不会出现两处筛选取交集不一致。
orphans 本身照旧返回,所以 PR 描述里「reported as an orphan rather than shown」这句仍然成立:被代链拒掉的那些不进 generations、不进选择器,但仍然列在 orphans 里报出来,不是静默丢掉。
| <footer class="page-footer"> | ||
| <p id="footer-source">数据来源:本机 MiniMax Code 会话记录,只读展示,不做任何写入。</p> | ||
| <p id="footer-binding"></p> | ||
| <p id="footer-truncated" hidden>文件过大,仅显示前 8 MB。</p> |
There was a problem hiding this comment.
这里写的是 8 MB,MAX_MESSAGE_BYTES 和 README 都是 64 MB。
There was a problem hiding this comment.
确认是我写错了,MAX_MESSAGE_BYTES 是 64 MB。
没有直接把它改成 64——那样两个数字还是各写各的,下次改常量又会漂。改成由 server 把它实际用的上限随 payload 一起发出来,页面用 formatBytes(session.maxMessageBytes) 渲染(index.html:4856),markup 里的初始文案不含任何数字。
| if (head.text === '') { | ||
| context.logger.warn('miniapp.trajectory.read_error', { reason: 'empty_messages' }); | ||
| } | ||
| if (head.read < head.size) truncated = true; |
There was a problem hiding this comment.
活跃文件的读取上限取的是 catalogBytes,文件正在追加写入、比 catalog 记录得长时,head.read < head.size 也成立,页面就会显示「文件过大」。复现时 688 字节的文件也出现了这个提示。建议只在真的触到 MAX_MESSAGE_BYTES 时才标记截断。
There was a problem hiding this comment.
复现得对,688 字节那个提示就是这么来的。两处一起改了:
- 读取上限不再取 catalog 记录的大小(那是快照,活跃文件一直在往后追加),只受
MAX_MESSAGE_BYTES约束,server.mjs:1837。readHead本身会按 stat 出来的真实大小收口,所以这个上限只管内存,不会再把新追加的内容切掉。 - 截断判定从
head.read < head.size换成truncatedByCap(head.size, limit)(server.mjs:218,调用点server.mjs:1850),只有真的顶到上限才算截断。
抽成单独一个函数是因为它一开始只写在一行里,没有能直接测到的地方。
| sendJson(response, 200, { | ||
| nodeVersion: process.version, | ||
| pid: process.pid, | ||
| ppid: process.ppid, | ||
| execArgv: process.execArgv, | ||
| argv, | ||
| envNames, | ||
| envSessionLike, | ||
| argvSessionLike: argv.filter((entry) => SESSION_ID.test(entry)), |
There was a problem hiding this comment.
README 写的是这个路由"no absolute path appears in the response",但 argv[0] 就是 node 可执行文件的绝对路径(Windows 上一般带用户名目录),execArgv、pid、ppid 也都原样返回了,README 里没有提到。页面也没有用到这个路由,建议删掉这几个字段,或者只保留 basename。
There was a problem hiding this comment.
确认违反了自己在 README 里的承诺。argv、execArgv、pid、ppid、envNames 全部删掉,连带那段扫环境变量找 session id 的死代码也一起删了(否则留着就是没人用的中间变量)。
现在只返回 nodeVersion、sqlite、cwdBasename、dataDirBasename、sessionRootExists。
整个路由没删:页面确实不用它,但它还能回答「运行时绑定活着吗」——sqlite 能不能打开、当前会话能不能解析出来。觉得该删我再删。
| async function openRuntimeStateDb(home) { | ||
| try { | ||
| const { DatabaseSync } = await import('node:sqlite'); | ||
| const dbPath = join(home, 'v2', 'sqlite', 'runtime-state.sqlite'); | ||
| if (!existsSync(dbPath)) return null; | ||
| return new DatabaseSync(dbPath, { readOnly: true }); | ||
| } catch { | ||
| return null; | ||
| } | ||
| } |
There was a problem hiding this comment.
每次调用都会新开一个连接,resolveCurrentConversation 和 readSessionTitles 用完都没有 close(),dispose() 也不会关。本地轮询时一直有 20–30 个句柄等 GC 回收,dispose 之后也还在。加个 try { ... } finally { db.close(); } 就可以。
There was a problem hiding this comment.
确认漏了。加了 closeQuietly(db)(server.mjs:620),两个函数都在 finally 里关(调用点 server.mjs:649、server.mjs:660)。
readSessionTitles 里 sessionIds.length === 0 的早退原本在 try 外面,也一并挪进去了,否则那条路径还是会漏。
The turn and call toolbar buttons were fold controls: they replaced rows with a collapsed placeholder carrying a one-line summary. They are now layer switches that remove rows from the render outright, leaving no placeholder behind and making the button the only way to bring them back. - turn governs the conversation layer (user, assistant) - call governs the tool layer (toolCall, toolResult) - the two layers do not overlap, so switching off one never takes part of the other with it - turn headers recompute from the rows that survive, so every count under a filter is a real post-filter count This drops the whole fold apparatus: collapsed/collapsedCalls state sets, reapplyFoldIntent, isCollapsed, isTurnFoldable, contentEntries, summarizeTurn, summarizeCalls, callBlocks, buildTurnPlan and the fold-summary styles. Button pressed state is now the visibility itself, so there is no derived state left to drift. Also in this change: - window controls: a sticky window bar offers load 200 earlier, jump back to the latest 200, and show everything. Show-all is a mode, not a big number, so incoming records do not knock it back down to 200 on the next poll. - scroll stability: the ledger restores a row anchor rather than an absolute pixel offset. The window shifts forward as records arrive, so a fixed pixel position drifts even though the scrollbar never moves. - sticky fix: .panel used overflow: hidden, which made it a scroll container and silently killed every position: sticky inside it. Switched to overflow: clip, which also revives the previously dead .turn-head sticky. - filtering: isFiltering and spanPassesFilter now agree entry for entry with visibleEntries. When they drifted, the overview bar drew spans for rows the list refused to render, which is what made filtering look incomplete. READMEs document the layer semantics and the new window controls.
`raw` is the untouched session message, and on a real conversation it is around two thirds of the payload. The page read it for exactly one record at a time — the inspector's raw tab, one clipboard copy, one handoff — so every poll transferred the weight of a session nobody was looking inside. The bulk payload no longer carries message bodies. A new /api/trajectory/raw serves one record out of the same cache entry that served the body, so opening a record costs one record rather than a session. Measured on a 5,614-record conversation: 27.0 MB before the split was even reachable, 10.8 MB after, and the session grew by 13 records in between. What did not move: `rawOmitted` stays on every row. Whether a message was too large to carry is a fact the list itself has to be able to state, and it is one boolean per row. The three readers are now strict. A fetch that fails says so in the panel or in the status line instead of quietly putting the normalized row where 原文 promised the original — those two are not distinguishable to a reader, which is exactly why the fallback is gone. A record whose message was too large to carry still falls back, since there is nothing better to show. handleTrajectory's session resolution is extracted into resolveTrajectoryRequest so the bulk payload and the raw fetch cannot disagree about which session a query string means. A 503 from the raw route means the server no longer holds the session; the page reloads it once and retries, capped at one retry so a permanently missing session cannot loop.
Hiding 轮次 left every thinking record on screen. Thinking arrives in the same model response as the answer it produced and shares its responseId, so it is part of what the turn was — and a list showing several hundred rows of it after the conversation was switched off reads as a filter that does not work rather than as one that does. CONVERSATION_KINDS now covers user, thinking and assistant. The button's own label said the old scope, so it reads 对话记录(用户、助手与思考)now. Nothing else moves: CALL_KINDS is untouched, the two layers stay disjoint, and a record whose message was too large to carry still reports rawOmitted on its row. verify-overview gains a direct assertion that no thinking record survives 轮次 being switched off; the fixture expectations for 轮次 off and for both off were derived by hand and corrected (both-off now leaves only the system row, since a3 was a thinking row that no longer survives).
Folding thinking into 轮次 answered the wrong question. Reading a conversation to find what was said, and reading a session to find out how the model got somewhere, want opposite settings for the same rows — so thinking is a layer of its own and the toolbar now reads 时长 / 轮次 / 思考 / 调用. CONVERSATION_KINDS goes back to user + assistant, THINKING_KINDS takes thinking, and a third flag, button, dom binding and click handler follow. All three layers stay disjoint and visibleEntries, spanPassesFilter and isFiltering all gained the same third check in the same place, so the overview strip and the list still agree span for span. The static aria-labels in the markup still said 点击折叠全部 — dead text that the first render overwrites, but it was the only place in the file still describing these as folds. They now state what the buttons actually do. verify-overview: 101 assertions, up from 96. The layer section was re-derived by hand — 轮次 off leaves ['a1','a3','a5','a6'] again, all three off leaves ['a1'], and the lit/out table now covers a third button, including a case that hides thinking while both other layers stay exactly where they were.
The toolbar's 轮次 / 思考 / 调用 and the 类型 chips were two independent filters over the same records. Five of the six chips mapped onto a button, so the page was asking one question twice and could answer it twice — which is what made the row look redundant. There is only one thing being filtered: which record kinds are visible. state.kinds is now the only place that lives, the buttons are presets over it (turn = user+assistant, thinking = thinking, call = toolCall+toolResult), and a button reads as lit only when every kind it covers is checked. Unchecking one chip puts its button out rather than leaving it claiming a half-on state the control has no vocabulary for. Gone: CONVERSATION_KINDS, THINKING_KINDS, CALL_KINDS, showConversation, showThinking, showCalls, and the three layer checks in visibleEntries and spanPassesFilter — hasKind was already doing exactly that work. isFiltering loses the three flags for the same reason. Net: fewer moving parts than before the buttons existed. The chips stay, because they reach subsets no button combination can: only user, only system, assistant without user. verify-overview asserts those explicitly so the chips cannot be "simplified" away later without something failing. setLayer splits into a pure applyLayerToKinds and the part that renders, matching how rePinWindow and turnTotalsFromEntries are already shaped here, so the rule a button applies is assertable without a DOM. Verified in the browser in both directions: unchecking the 思考 chip put that button at aria-pressed=false while the other two stayed lit, and clicking the button put the chip back. verify-overview 113 assertions, up from 101.
Unifying the buttons onto the chip set left one case the buttons could not say out loud. Hiding 助手 in the chips while leaving 用户 checked puts 轮次 out — "lit" can only mean "all of it is on" — but the reader did not hide the conversation, they narrowed it, and the button's plain 已隐藏 claimed otherwise. It was also the one case where clicking the button silently overrode a choice the reader had just made. A button now has three states: lit when every kind it covers is checked, 当前已隐藏 when none are, and 当前部分显示 · 点击全部显示 when some are. That wording says both that the layer is narrowed rather than gone, and that clicking fills it in. 思考 covers a single kind and so can never be partial. layerPartial is pure and sits beside layerSelected, and verify-overview now pins the behaviour: a partial layer is neither lit nor fully cleared, says the right thing, leaves the other layers alone, and clicking it restores only the kinds that were missing. Verified in the browser: unchecking only 助手 left 轮次 at aria-pressed=false reading 当前部分显示 · 点击全部显示 with the other two buttons untouched, and clicking it put the 助手 chip back on. verify-overview 121 assertions, up from 113.
The three buttons above the list were presets over the type chips, which made them a second way to ask the same question. Any preset a reader outgrows ends up contradicting the chips it was meant to summarise: untick 助手 while leaving 用户 ticked and the turn button goes dark claiming the conversation is hidden, when in fact it was narrowed. There was a third button state written purely to explain that contradiction. There is now one control and one answer. state.kinds is written from three places - the chips, 全部, 清除筛选 - and client-check asserts that count plus the absence of the layer machinery, so this cannot quietly come back. 时长 moves down into the filter bar next to the chips: it changes how the overview draws the clock, not which records exist, so it reads as a view control sitting beside the filter rather than as another filter above it. The removed 32px toolbar row was carrying the sticky offsets, so .overview now pins at 0 and .window-bar at 51px (50 plot + 1 border). Without that the window bar would slide under the strip.
Two defects with the same shape: a control or a label that could not be acted on the way it appeared to be. A compactionSummary was filed as kind 'system' with title 'compaction', so the client had to special-case the title in five places to keep calling it 压缩. The filter reads kind alone, which made a context checkpoint impossible to keep or drop on its own - unticking 系统 took every checkpoint with it. It is now a seventh kind, which deletes all five special cases and gives it a chip. The axis button was labelled 时长 and sat with the filters at the top of the page. That reads as a filter on duration; it changes neither the rows nor any count. It is renamed 总览轴, its label now names the strip and both axes, and it moves into the window bar directly under the strip it redraws. Putting it there meant the window bar could not keep rebuilding its nodes: a render happens every poll, so a click landing between the clear and the re-append went to a detached button. It is built once now and updated after.
The ledger builds a real element per row — there is no virtual scrolling — so
rendering the whole session at once is not an option to offer by default. On a
9,900-record session that is roughly 57,000 elements rebuilt every 2.5s poll,
which was measured to freeze the page. The page now opens on the newest 200 and
the window bar switches to everything on request.
- The button names the action pressing it performs, not the state it is in:
「显示全部」 at the default, 「显示最新 200 条」 once everything is showing.
The count beside it reports what is on screen; the two answer different
questions and so cannot contradict each other. `aria-pressed` carries
whether the full view is on and `aria-label` restates it as a sentence.
- 「向上加载更早的 200 条」 widens the window, and a poll no longer undoes it.
`rePinWindow()` used to reset the limit unconditionally, so the button never
actually worked — every 2.5s poll took it straight back to 200. It now
leaves a widened window alone and only re-pins a window that is exactly the
default size, which is what "follow the tail" means.
- The button does not scroll the page away from the reader's place.
- `.btn[hidden]` is now honoured. `display: inline-flex` was overriding the UA
rule, so a hidden button stayed on screen with no label in it. Six sibling
selectors already carried this patch; `.btn` was the one that did not.
- `DEFAULT_WINDOW_MODE` is a single constant that every reset reads, rather
than eight copies of the literal that could disagree with each other.
Six review comments:
- An empty `records` array was read as "this whole turn is empty"; it is per
message, so the test now looks at this message's `segmentOrder`.
- `generations` listed the whole catalog, so a generation the lineage walk had
rejected — a snapshot left behind by a fork — was counted into the
session-wide total and offered in the picker, where selecting it resolved to
an empty view. The payload now reports the chain, and still discloses the
rejected ones through `orphans`.
- The footer claimed a cut at 8 MB while the server reads 64 MB. The cap is
sent with the payload and rendered from it, so the two cannot drift again.
- The active file's read was bounded by the catalog's recorded size. That size
is a snapshot and the file keeps being appended to afterwards, so an actively
written session could stop the read short of data that was really there and
then report itself as "file too large". The read is bounded by our own cap,
which is the only thing that can actually truncate, and `truncatedByCap()`
is now the single place that decides it.
- `/api/runtime` returned `argv` (whose first element is the node path, and on
Windows a username), `execArgv`, `pid` and `ppid`. The page never reads that
route; it now reports only the four facts it can state safely.
- `resolveCurrentConversation` and `readSessionTitles` opened a SQLite handle
per session and never closed it, so a long session picker leaked a handle per
entry. Both close in a `finally`, including on the early return.
A checkpoint was filed under "<turnId>:compaction" while every row after it carried the bare id, so keying turns on the whole string opened a turn for the checkpoint alone. Turns join the list the moment they are first seen, and the interrupted turn kept collecting records after the checkpoint had been seen — so the checkpoint was emitted after the entire turn it belonged to. On a 10,500-record session that showed up as 12 turns of exactly one row each, a numbering that ran backwards (#1037 then #970), and timestamps stepping back up to 20 minutes. Keying the checkpoint on the bare id puts it back where it happened. Measured on the live session: turns with a hole in their index range 12 -> 0, timestamp inversions caused by a checkpoint 12 -> 0, and the ones that remain (17, all a steer after the tool result it interrupted) are left alone on purpose — a steer carries the time the reader typed, and sorting it back into place would put it before the calls it was answering. A turn that *contains* a checkpoint is still judged on what it did. "Contains one" used to be the same as "is one" back when a checkpoint was a turn of its own; keeping that would relabel every compacted turn and drop all of them out of the prompted count. Only a turn holding nothing but its checkpoint is one. The checkpoint is drawn as a full-width band rather than as another row. It is still one filterable record — the 压缩 chip counts and hides it like any other — but the rows above and below it are no longer the same conversation as far as the model is concerned, and one line says so instead of leaving a gap to infer from. That band's number now comes from the record count, because the turn count is zero for every session that was ever compacted. Both counters were also being kept honest about what they count: a record's `turnId` now names the turn it is actually in, since the file's suffixed id names no turn at all.
Reported as "#1085 appears more than once". No index was duplicated: the payload holds 11030 records and 11030 distinct indexes, and the ledger builds one row per entry. The number was being cut off on screen. The first grid track was 42px, which holds "#" plus four digits — the whole of every session until one passed 10,000 records. Past that the fifth digit left the track, and a grid item is `overflow: visible` by default, so it spilled into the 76px kind column and was painted over by the opaque pill. #10850..#10859 all read #1085; ten consecutive rows showed the reader one number. Nothing errored, and unlike a wrong number this one cannot be checked — it looks exactly like a real index. Three grids render a record index and all three were 42px; the max-width: 767px override narrowed it to 28px, too small even for four digits, which is what made the phone layout the worst case. They now share one `--traj-index-col: 72px`. 72px is derived, not chosen: the old track held five characters and would not hold six, so one character is at most 42 / 5 = 8.4px, and a label is "#" plus at most seven digits — 8 x 8.4 = 67.2px. A session would have to reach 10,000,000 records before this bites again. .row-index also gets `text-overflow: ellipsis` so that if a column is ever too narrow again the row says it could not show the number in full, instead of quietly displaying a different valid one. Verified in the browser at 72px: .row-index measures width 72 on both the plain row (#11023, x 31-103) and the compaction band (#11034, x 33-105), with the kind pill starting at x 109 and x 113 respectively, and consecutive indexes #11049..#11069 render complete and distinct.
Adds a 时间区间 row under the search box — 起点 / 终点 boxes plus 更新到最新 and 清空区间. Empty by default, and the table, the window count and the KPI strip all read the same one filter. - Input is a session-relative offset (`1:30`, `1:30:00`, a bare number reads as minutes) because the ledger already prints `+893 分 21 秒`; wall-clock lines up with nothing else on the page. The hint always echoes the parsed value back, so what the filter is actually using is never a guess. - KPI is recomputed client-side from the same `state.flat` records the table reads, and the probe proves it equals `payload.stats` field for field across all 15 fields on the real 11,812-record payload. Header and table cannot drift onto two definitions. - A record matches on its own offset; a turn matches when any record inside it matches, and is then counted whole. `messagesAll` stays 全场 so the comparison label still means something. - The range is a snapshot, not a subscription. `rangeCutoff` pins the highest record index when the range is set, so new records cannot append themselves, and 更新到最新 is what moves it. Measured: 共 11,679 → 11,685 while 已显示 654 held still. - Setting a range calls `resetWindowMode()` like every other filter. Forcing 「全部」 was measured and dropped: a range starting at 20:00 covers the whole session, and 全部 then rendered ~11,000 rows and took the document to 414,000px, where one query took 19s. 不自动追加 does not depend on it — the cutoff holds the range on its own. - A box the page cannot read is never written into state. `NaN` fails the `isCount` test behind `hasTimeRange`, so storing it would switch the whole range off mid-keystroke and quietly hand back the entire session. The hint reads the boxes rather than state, which is what makes the 「格式无法识别」 message reachable at all. - 清空区间 hides itself while no range is set, and 清空 clears the range too.
… offset
The boxes took a session-relative offset, which is what the ledger prints on every row — and
which stops being answerable once a session runs past midnight. This one crosses 2026-10-11 at
offset 12:14:29: `12:00:00` is 2026-10-10 23:45:31 and `13:00:00` is 2026-10-11 00:45:31, so an
hour of offset covers three quarters of an hour of clock with midnight in the middle and nothing
on the page saying so. A reader asking for 「23:50 到 00:20」 had no way to ask it.
So a box may now hold either clock, and the text decides which.
- The two grammars cannot be confused because they share no separator: an offset is digits and
colons (`1:30`, `1:30:00`), a date is digits and `-` or `/`. `hasRangeDate` reads that off
the text before any value is computed, so the filter, the label and the hint cannot come out
disagreeing about what was typed.
- The conversion is a subtraction and nothing more. Every record already carries `timestamp`
and `relativeMs`, and `session.startedAt` is the origin those offsets were measured from, so
no second clock is introduced anywhere. Checked against the live payload: a date range over
23:50–00:20 selects exactly the 545 records whose `timestamp` falls in it, and the
equivalent offset range `12:04:28 – 12:34:28` selects the same 545.
- The year is required rather than inferred. Guessing means deciding what `01-01` means in a
session that starts in December, and a range that lands on the wrong year is worse than one
the reader has to finish typing.
- The date is assembled from parts and round-tripped, never handed to `Date` as a string:
`new Date('2026-10-10')` is UTC, and without the round trip `2026-02-30` rolls silently into
March and returns a range nobody asked for.
- A bound before the session began stays negative instead of being clamped, and now prints its
sign — `Math.abs` on its own made `-2:45:00` read as `2:45:00`, which in the one place the
sign carries the whole message is the wrong answer rather than a formatting nit.
- Each end of the label comes back in the clock its own box was written in. An offset beside a
date is a mismatch the reader has to decode before trusting it.
- A session whose records carry no timestamp says so, instead of reporting 「格式无法识别」 and
sending the reader hunting for a typo that is not there.
- The boxes were a bare 88px, which held `2026-10-10` and scrolled the time out of sight. The
value stayed correct and the hint echoed it back, but the field could not be checked by
looking at it — the same defect as the 序号 column. Width is now 18ch, derived from the
rendered box: 15ch came out at 113px and held 15 character slots, and sixteen glyphs plus
the caret is 17.
Probes: 162 assertions in verify-time-range.mjs, and 55 + 12 mutations in the two runners, each
caught by the guard that owns the defect.
…et grammar The boxes took a session-relative offset, on the argument that `+893 分 21 秒` is what every ledger row prints. They are now two `datetime-local` pickers with four quick buttons beside them, and the offset grammar is gone. The argument did not survive contact with a session that crosses midnight. This one crosses 2026-10-11 at offset 12:14:29, so `12:00:00` is 2026-10-10 23:45:31 and `13:00:00` is 2026-10-11 00:45:31 — an hour of offset covering three quarters of an hour of clock, with the midnight in between and nothing on the page saying so. It also did not survive switching sessions, where the same number means a different moment. A picker is the honest control for a moment in time, and one grammar per box is one less thing the reader has to work out. The quick picks are the four windows this page can answer for any session: - 最近 15 分钟 / 最近 1 小时 — the trailing history a reader wants when something just happened - 今天 / 昨天 — the two wall-clock days a session that runs overnight is split across 清空区间 already means the whole session, so there is deliberately no 全程 button beside them. Days are stepped through the `Date` constructor rather than by subtracting 86,400,000, which lands on the wrong day across a daylight-saving boundary, and a pick whose window opens before the session is clamped to the session start — "the last fifteen minutes" of a five-minute-old session is its whole life, not a negative range. Two things the picker could not do on its own: - A window that is well formed and lies outside the session now says 「这段时间里没有记录」 and names where the session actually starts. That is exactly what 昨天 does on a session that began this morning, and without it an empty table reads as 「这个会话没有内容」. - The pickers are bounded by the span the session covers, re-decided on every render because the session grows during the first poll. The controls also had to stop being the odd one out: `datetime-local` was missing from the shared input rule, so it arrived at its own intrinsic 22px height with the user agent's own border. And it is now sized by its own content rather than by a number in the stylesheet — the same defect as the 序号 column, committed twice: 18ch clipped the year off the text box that preceded it, and 20ch clipped the minutes behind the calendar icon on the picker. The browser lays a `datetime-local` out from the localized date it has to draw, so a width in the stylesheet is a guess about something the user agent already knows, and it cannot go stale the way a measured constant does. Probes: 148 assertions in verify-time-range.mjs, and 66 + 11 mutations across the two runners, each caught by the guard that owns the defect.
An interruption is the thing a reader of this page most wants to find, and until now nothing pointed at one. 失败 was a pill on a row and 压缩 was a band, so those two were findable by eye; a turn that simply took a prompt and never answered had no marking anywhere, which is the shape an interrupted turn leaves behind. So 异常 is a filter, not a kind. A kind is a label the runtime stamps on a row; 异常 is a question asked across the whole ledger — 「这次运行在哪儿出的问题」 — and its answer lands on three unrelated kinds at once. Filing it as a tenth kind would have put it in the type chips, where it reads as one more category instead of the second axis it actually is. Each rule is one predicate plus the label the reader sees when it fires: - toolError 工具失败 — a tool result the runtime marked as an error - compaction 上下文压缩 — the checkpoint written when the history underneath a turn is rewritten - noReply 无答复 — a turn that took a prompt and never produced a reply Adding a rule is adding an entry to that table. The buttons are built from it rather than written into the markup, so a rule has exactly one name in exactly one place. The kind chips already had this defect once: the label lived in the markup and the accessible name was written beside it, so a rename left a button showing one word and announcing another, and nothing in the guards could see it because both halves were individually correct. The turn still running is excluded from 无答复. It is open because it is going, which is not the same thing as broken, and a rule that re-flagged the newest thing on screen on every 2.5s poll would be a filter nobody could leave on. The counts on the buttons and the rows in the list come from one predicate. anomalyCounts() walks the flat list to fill the buttons, and while that walk was its own copy a button could promise 246 while the list showed 12 — the same 「筛选了但什么都没变」 this page keeps guarding against, moved up one level. Both go through passesBaseFilter() now. The overview strip is handed a bare record and 无答复 is a property of the turn, so the owning turn comes back out of a turn index; without that lookup the strip and the ledger directly above it answer different questions about the same row. 无答复 is marked on the turn head and not on each row. It is one thing that happened to a whole turn, and turn 82 alone would have carried the tag 352 times — at which point it stops being a signal and becomes wallpaper. Measured on the live session: 94 turns, 13,132 records, 1,416 anomalous — 251 tool failures, 15 checkpoints, and 1,184 records spread across 10 turns that never answered. Probes: 98 assertions in verify-anomaly.mjs against the running session, with the expected counts taken straight off the payload by code that shares nothing with the page, and 27 mutations each caught by the guard that owns the defect. Two guards in verify-time-range.mjs that pinned the old shape of visibleEntries() and clearFilters() were rewritten to pin the invariants instead of the layout, and the reset-call inventory moved from ten to eleven for the new filter.
Tracing an interruption back to the raw session file turned up a shape the 异常 filter could not
see, and the reason is worth writing down because the rule set was built on the assumption that
every anomaly is a *present* signal — a flag, a kind, a missing turn.
The turn the reader watched finish without an answer was closed like this:
"role": "assistant", "content": [], "stopReason": "stop",
"usage": { "output": 78, ... }, "responseId": "071a35d49ed..."
The call happened. There is a response id and an output token count. It stopped normally, not
aborted, not timed out. It just came back with nothing in it. The turn therefore has a reply
record, `closed` is true, and 无答复 — which tests `closed === false` — has nothing to fire on.
The 答复 chip counted it the whole time. From the ledger's side it was a reply; from the reader's
side nothing was ever displayed.
So the new rule tests an absence:
record: function (record) {
return !!record && record.kind === 'reply' && isNil(record.text);
}
It keeps its kind on purpose. The record genuinely is a reply, the 答复 chip counts it and the
reply filter keeps it — it is an anomaly *as well as* a reply, not instead of one. `text` is the
field to read rather than `textTruncated`: a clipped reply still has a body, however short, and
whitespace is a body a reader can see.
The row needs to say so too, or the filter shows rows that look like ordinary answers. It takes
over the duration cell the way 失败 already does, so the five-column grid does not grow, and the
name comes from the rule table rather than a second copy written at the call site — the kind
chips already paid for having the visible word and the announced word in different places.
Measured on the live session: 99 turns, 13,358 records, 1,422 anomalous — 255 tool failures, 15
checkpoints, 1,184 records across 10 turns that never answered, and 2 replies that delivered
nothing. Both empty replies sit in closed turns, which is checked against the session rather
than assumed, because that overlap is the whole reason this rule exists.
Probes: 127 assertions in verify-anomaly.mjs and 37 mutations, each caught by the guard that
owns the defect. Two of the new assertions are about what the rule must NOT do — a checkpoint
and a bodyless system record are not empty answers, or the union would come out smaller than
its parts.
MyPrototypeWhat
left a comment
There was a problem hiding this comment.
谢谢这么快就逐条改完!上一轮的 5 处我都用同一套合成数据和 headless Chrome 重新跑过:空内容的助手消息现在有对应的行,孤儿快照不再出现在选择器和「全程」里,截断提示改成 64 MB 而且追加写入时不再误报,/api/runtime 不再返回 argv/pid,SQLite 句柄在轮询和 dispose 之后都是 0。新加的时间区间、异常筛选、按需原文(包括缓存过期后自动重试)我也试了,都工作正常。
需要修改
-
README 有几处没跟上代码:
- Diagnostics 段落(README.md:94、96,zh 84、86)还在说
/api/runtime会返回环境变量名,并检查 env/argv 里有没有会话 id,以及"reports 13 environment variable names",这些已经删掉了; - README.md:25(zh 25)说"The type chips are the only filter",但这一轮新加了「时间区间」和「异常」两组筛选,两份 README 里都还没提到;
- README.md:72(zh 62)说用 catalog 长度给读取设上界,现在活跃文件只受
MAX_MESSAGE_BYTES限制,不读半行靠的是readHead截到最后一个换行。
麻烦改成和现在的行为一致;PR 描述如果能顺手补上这些新功能就更好了。
- Diagnostics 段落(README.md:94、96,zh 84、86)还在说
建议(可选)
/api/runtime里的probe(server.mjs:1642-1647)只在查询成功时close()。表结构不匹配时会留下句柄等 GC,换成finally { closeQuietly(probe) }就行。- README.md:62 "…says so where it shows.One model call…" 两段连在了一起;README.md:113 "offers that choice backers four explicit controls" 这句读不通。
- 上一轮的三条可选建议(缺 catalog 的会话打开会 500、
history_relative_dir的包含检查、MMC_TRAJECTORY_ROOT的措辞)还在,看你方便。
The second review round confirmed the five code fixes from the first (segmentOrder, orphan filtering, the 64 MB cap, the runtime route, the SQLite handles) and re-ran them, but the READMEs had been left describing the app as it was before the time range and anomaly filters landed. Bring them back in line, and close the open suggestions while the review is still warm. README, both languages: - "/api/runtime" still advertised environment variable names and an argv scan. Both were removed in the previous round; say what the route returns now and why the scan is gone rather than quietly dropping the sentence. - "The type chips are the only filter" was true when written and is not now. The time range and the anomaly filter cut across kinds rather than over them, so state that instead of leaving the reader with one axis in a UI that has three. - The read was described as bounded by the length history-catalog.json recorded. It is bounded by MAX_MESSAGE_BYTES, and readHead clamps to the last complete line. The cap is sent with the payload and rendered from there, so the README should stop carrying its own copy of the number. - MMC_TRAJECTORY_ROOT replaces the data directory; v2/sessions and v2/sqlite are still resolved beneath it. - Two paragraphs had run together at README.md:62, and the sentence offering "four explicit controls" had lost its verb. server.mjs, three suggestions that were still open: - /api/runtime closed its probe handle inside the try. A schema mismatch is precisely the case this route exists to report, and it is precisely the case where prepare() throws -- so the handle waited on the GC every time. - history_relative_dir was joined onto the sessions root without the containment check resolveArtifactPath gives catalog file names. It is a database column rather than a name this app chose; it gets the same check. - A session with no history-catalog.json -- one that has never compacted -- used to answer artifact_path_rejected and a 500, because artifactByFile was empty even though the active file was real. probe() already synthesised an artifact for exactly this case; reuse that shape instead. npm run check passes (validate clean for this plugin, 55/55 tests) and the existing probes are unchanged: verify-pr15-fixes 26, verify-generations 108, verify-compaction-turn 22, verify-time-range, verify-anomaly 127, verify-overview 108, client-check 168, plus the response-rail, window-widen, call-fold, block-order and prompt-fix probes.
MyPrototypeWhat
left a comment
There was a problem hiding this comment.
感谢更新,这一轮我又在本地对 871c5d5 逐项复现了一遍:
- README 的 Diagnostics 部分(中英文)现在和
/api/runtime实际返回的字段一致,筛选说明也补上了时间区间和异常筛选;读取上限的说法改为以MAX_MESSAGE_BYTES为准,和代码行为吻合。 /api/runtime的 probe 放进了finally:表结构不匹配时连续请求 20 次,SQLite 句柄数仍然是 0。- 没有
history-catalog.json的会话现在能正常打开(HTTP 和页面上都验证过);history_relative_dir越界时会被拒绝,并回落到按 mtime 判定;MMC_TRAJECTORY_ROOT的说法也和实现对上了。 - 之前的 5 个必改项都没有回归,Node 22 下
npm run check全部通过。
一个小建议(可选):PR 描述里提到「如果规则名被写成字面量,会有 guard 报错」,代码注释里也提到 offline probe,但仓库里没找到对应的脚本。可以考虑把它提交进来(放在 miniapp/ 之外),或者在描述里注明是本地检查。
整体没有问题,LGTM。CI 还需要维护者批准后才会运行。
|
Closing the loop on the last note in this thread — the rule-name guard was not Two notes on what is deliberately not in it, both recorded in
Thanks again for the three rounds — the six fixes in #15 are the reason those guards exist at all. |
What
Adds Session Trajectory (会话轨迹), a Mini App that renders the model trajectory of a
MiniMax Code conversation: messages, reasoning, tool calls and their results, Token usage,
and per-turn timing, grouped by
turn_id.Why it exists
A Mini App cannot tell which conversation opened its page — the runtime context has no
session id, the client bridge exposes only
miniapp.message.append, and the page URL carriesno parameter. This app resolves the active conversation from the local runtime database,
always states in the page footer how it decided, and reports ambiguity instead of silently
picking when more than one conversation holds a live turn lease.
Data access
<dataDir>/v2/sessions/**/messages.jsonl,snapshots/*.jsonl,manifest.json,history-catalog.json;<dataDir>/v2/sqlite/runtime-state.sqlite(read-only)This relies on an unspecified internal layout (
v2/sessionsstructure and the runtimedatabase schema). A client update can break it. The README documents this, and there is a
capability request for a supported way to bind a Mini App to its session:
MiniApp capability wishlist #9.
Context generations
A compaction does not trim a session in place — the runtime rotates the previous
messages.jsonlintosnapshots/, opens a new generation, and starts a fresh active file. Sothe ledger shows the whole conversation, not just the tail. The app walks the chain the way
the runtime does: it opens the active file, reads the generation that file declares in its own
history_artifact, and steps back one generation at a time through the parent each file names.A snapshot whose parent generation is not exactly one lower ends the chain instead of being
stitched on; a snapshot the chain never reaches is reported as an orphan rather than shown.
A 上下文代 picker appears when a session has more than one generation and narrows the
ledger to a single slice. Every record carries the generation it came from.
Filtering
Three controls, on two axes — and the distinction is the point.
Over kinds. The eight kind buttons, 全部 and 清除筛选 write one array. Nothing else touches
it. That is why the row of layer presets that used to sit above the list was removed: a preset a
reader outgrows turns into a button that contradicts the chips it was supposed to summarise.
Across kinds. Two controls narrow the ledger without rewriting that array:
时间区间 — two wall-clock pickers, plus
15m/1h/today/yesterday. A relativeoffset grammar was built first and then removed: an offset silently means a different interval
as the session ages, and on a session that crosses midnight "the last hour" is ambiguous.
异常 — rules that pick out records which are wrong, not records of a particular kind:
工具失败上下文压缩无答复空答复Every rule name lives in one table on the client. The buttons are generated from it, the row and
turn-header markers read their labels from it, and a guard fails if a rule name is ever written
as a literal instead.
The empty answer is worth calling out because it is invisible by construction: the record is an
ordinary
replywhose content is empty, so nothing looks broken unless you count replies thatproduced no text. It is deliberately both a reply and an anomaly — unticking 答复 hides it,
and the anomaly filter finds it.
Test environment
Verified manually in the client: publishing and startup, live polling during an active turn,
session switching, type filtering, time-range filtering (including across midnight), the anomaly
filter and its interaction with the type chips, on-demand record bodies, error states. The session
used for verification has been compacted twice and spans three generations — 2,000+ messages and
11.6 MB across the whole lineage, loaded without truncation. Session resolution was cross-checked
against the conversation title shown in the client and against the session's own manifest.
Not verified: macOS and Linux, the light theme, and compatibility with non-official or
source builds.
Notes for reviewers
node:built-ins; zeronode_modules.miniapp/node/miniapp-api.tsis a type declaration and is never imported at runtime.—; nothing is inferred or filled in.userare split off via the recordedcanonicalTextRangeand shown as系统, so the ledger never presents injected context assomething the user typed.
opened a context generation is a checkpoint, and a reply that arrived through a question card
never became a typed message, so counting only what was typed would not reconcile with the
total.
304with no body; a validator is per generation, so a narrowedview cannot be satisfied by the stitched one.
MAX_MESSAGE_BYTES, not the lengthhistory-catalog.jsonrecorded: the activemessages.jsonlkeeps being appended after the catalog is written. The cap in force is sentwith the payload and rendered from there, so the number is written once.
history-catalog.json— one that has never compacted — is read from itsactive file alone rather than refused.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.