Repository navigation
[Due for payment 2026-10-15] [$175] LHN sometimes doesn't show unread chats until the chat is opened #102973
Description
Activity
- addedDailyKSv2KSv2Needs ReproductionReproducible steps neededReproducible steps neededBugSomething is broken. Auto assigns a BugZero manager.Something is broken. Auto assigns a BugZero manager.
on Oct 5, 2026 Proposal
Confidence: medium. The code and the PR that introduced the change support this. I couldn't check production logs in this run, so it isn't confirmed for a specific user yet.
What is the root cause of that problem?
The LHN is fine. On web, the new-message update never reaches the client.
Since [No QA] fix: skip reconnectApp on pusher socket blips (merged 2026-08-29), web skips
ReconnectAppwhen the Pusher user channel resubscribes, unless pusher-js reached theunavailablestate (PusherUtils.ts:58-67,Pusher/index.ts:115-118).- A socket dies silently, for example after sleep, a Wi-Fi change, or a throttled tab. pusher-js only notices this after its activity and pong timeouts.
- It then reconnects in under 10 seconds, so it never reaches
unavailable. The resubscribe skips the sync. - Pusher doesn't replay events. Messages sent while the socket was dead are lost.
- Gap detection only runs when the next server payload with
previousUpdateIDarrives (applyOnyxUpdatesReliably.ts:50-51). Opening the chat sendsOpenReport, which triggers it. That's why the unread state appears only after the user opens the chat.
PR #99084 calls this out in its own test step 10: "Verify the message does not appear yet. This is the accepted tradeoff… an idle tab stays stale until one arrives."
What changes do you think we should make in order to solve the problem?
- In
src/libs/Pusher/index.ts, also mark an outage when the socket drops after an activity or pong timeout (a silent death). Optionally, also mark one when the time since the last inbound socket message is longer than a few seconds. Short blips still skipReconnectApp, and silent deaths sync again. - Add a unit test to
tests/unit/PusherSubscribeWebTest.ts. It should cover a pong-timeout drop followed by a fast reconnect, and check thatReconnectAppfires. - For the telemetry the issue asks for, add a
Log.infowhen gap detection runs. It should include the gap size and the time since the last "Skipping reconnect" line. Then we can measure how often skipped syncs leave clients stale.
What alternative solutions did you explore? (Optional)
- Revert [No QA] fix: skip reconnectApp on pusher socket blips #99084. This is simple but brings back about 54k extra
ReconnectAppcalls every 2 hours, per the PR. - An LHN rendering bug. I traced the code from
isUnreadto the LHN row and found no memo or cache that skips a changed report. Opening the chat repairs the data itself, which points to missing data, not a render bug.
Investigation details
- Web only: native's
claimOutageSyncalways returns true (index.native.ts:481-483). - Timing: [No QA] fix: skip reconnectApp on pusher socket blips #99084 shipped about 2026-08-31, which matches "past few days or weeks."
- Related changes: [No QA] stop the pong watchdog from reconnecting Pusher (2026-08-19) removed the earlier safety net for [$250] False "offline" indicator after machine restart — stale Pusher websocket not re-established. Lower the Pusher activity and pong timeouts from the SDK defaults (2026-09-14) means silent deaths now reconnect fast and always take the skip path.
- How the LHN decides a row is unread:
isUnreadcompareslastReadTimewithlastVisibleActionCreatedon the report (ReportUtils.ts:10016-10046,SidebarUtils.ts:844). If the report update never arrives, the row stays read. - Less likely cause: the chat was auto-marked read by another mounted screen, tab, or device (
useMarkAsRead.ts:199-241). This fits less well, because the unread state does appear after the chat is opened. - To confirm with logs: for an affected account, look for
[PusherUtils] Skipping reconnect, socket recovered without going unavailable, then noHandled multipleEventsline for the missed message, then a gap-detected line when the user opens the chat. - Questions for reporters: When this happens, does the LHN preview text or row order also not update? Is there no notification? Are you using the Unread tab or All?
Possibly related issues:
- Marking a chat as unread doesn't stick — client immediately re-reads it, and other clients never see the change (open). Its bug 2 is a client that never receives updates.
- [$250] Messages not highlighted as unread in LHN on web / real-time message updates not reflected (closed, couldn't reproduce). Same symptom, before [No QA] fix: skip reconnectApp on pusher socket blips #99084.
Next Steps for Contributor+ team:
To accept:@MelvinBot implement [this](https://github.com/Expensify/App/issues/102973)to create a draft PR.
To refine:@MelvinBot <your feedback>
To reject: Explain why you are rejecting Melvin's proposal.
Reacted by Viktoryia Kliushun- addedExternalAdded to denote the issue can be worked on by a contributorAdded to denote the issue can be worked on by a contributor
on Oct 5, 2026 Triggered auto assignment to Contributor-plus team member for initial proposal review - @dmkt9 (
External)- changed the title
[-]LHN sometimes doesn't show unread chats until the chat is opened[/-][+][$175] LHN sometimes doesn't show unread chats until the chat is opened[/+]on Oct 5, 2026 Job added to Upwork: https://www.upwork.com/jobs/~022107043087024337291
📣 @dmkt9 We're missing your Upwork ID to automatically send you an offer for the Reviewer role.
Once you apply to the Upwork job, your Upwork ID will be stored and you will be automatically hired for future jobs!Proposal
Please re-state the problem that we are trying to solve in this issue.
On Web, unread chat messages do not highlight as unread in the Left Hand Navigation (LHN) until the user clicks and opens the chat report.
What is the root cause of that problem?
In
src/libs/PusherUtils.ts(lines 60–67), when Pusher re-subscribes to the private user channel, it callsPusher.claimOutageSync():unregisterPrivateUserChannelResubscribe = Pusher.onChannelResubscribe(getUserChannelName(accountID), () => { if (!Pusher.claimOutageSync()) { Log.info('[PusherUtils] Skipping reconnect, socket recovered without going unavailable'); return; } Log.info('[PusherUtils] Pusher re-subscribed to private user channel, triggering reconnect'); reconnect(); });
In
src/libs/Pusher/index.ts(line 116),hasUnclaimedOutageis only set if the socket reaches theunavailablestate:hasUnclaimedOutage ||= states.current === 'unavailable';
When a socket connection drops silently (e.g. laptop sleep, Wi-Fi network switch, browser throttling idle tabs, or activity/pong timeout) and reconnects in under 10 seconds,
pusher-jsnever transitions intounavailable.Consequently:
claimOutageSync()returnsfalse, skippingreconnect()(ReconnectApp).- Pusher does not replay missed real-time events that occurred while the socket was disconnected.
- The report's Onyx state is never updated with the new message until the user navigates into the chat, which triggers
OpenReportand fetches the latest report data.
What changes do you think we should make in order to solve the problem?
- In
src/libs/Pusher/index.ts:- Track when a socket disconnect occurs due to transport error or connection timeout (such as activity/pong watchdog drop) by setting
hasUnclaimedOutage = true. - Ensure that any resubscription that occurs after an unnoticed connection break properly triggers
ReconnectAppviaclaimOutageSync(), while keeping immediate blips without missed events optimized.
- Track when a socket disconnect occurs due to transport error or connection timeout (such as activity/pong watchdog drop) by setting
- In
src/libs/actions/applyOnyxUpdatesReliably.ts:- Add telemetry logging in gap detection to monitor gap size and delta time since the last skipped reconnect to ensure zero dropped syncs.
- Unit test:
- Add a test case in
tests/unit/PusherSubscribeWebTest.tsasserting that an activity/pong timeout drop followed by a fast reconnect correctly marks an outage and triggersReconnectApp.
- Add a test case in
What alternative solutions did you explore? (Optional)
- Reverting PR [No QA] fix: skip reconnectApp on pusher socket blips #99084: This would trigger excessive
ReconnectAppcalls on transient socket blips (~54k extra calls). Targeting actual timeout/drop disconnects rather than relying solely onunavailablepreserves efficiency while eliminating stale state. - Client-side polling in LHN: Unnecessary and contradicts Onyx push-driven architecture.
Contributor details
Your Expensify account email: camesenin@gmail.com
Upwork Profile Link: https://www.upwork.com/freelancers/~0170d6b02e9440e6045 remaining items
- addedReviewingHas a PR in reviewHas a PR in reviewWeeklyKSv2KSv2and removedDailyKSv2KSv2
on Oct 6, 2026 Updates:
- Today I was able to reproduce one of the scenarios mentioned in slack when unread message isn't display (when new message arrives in the thread):
thread-case.mp4
- Prepared PR with the fix: Recheck LHN reports when their derived attributes change #103147, applied some C+ feedback, waiting for another round of review
PR merged.
- addedDailyKSv2KSv2and removedReviewingHas a PR in reviewHas a PR in reviewWeeklyKSv2KSv2
on Oct 8, 2026 - changed the title
[-][$175] LHN sometimes doesn't show unread chats until the chat is opened[/-][+][Due for payment 2026-10-15] [$175] LHN sometimes doesn't show unread chats until the chat is opened[/+]on Oct 8, 2026 The solution for this issue has been 🚀 deployed to production 🚀 in version 9.5.5-2 and is now subject to a 7-day regression period 📆. Here is the list of pull requests that resolve this issue:
If no regressions arise, payment will be issued on 2026-10-15. 🎊
The following checklist (instructions) will need to be completed before the issue can be closed. Please copy/paste the Contributor+ Checklist from here into a new comment on this GH and complete it. If you have the K2 extension, you can simply click: [this button]. If no checklist is needed for this issue, you can click: [no checklist button]
Contributor+ Checklist:
-
[Contributor] The offending PR and associated issue have been commented on, pointing out the bug it caused and why, so the author and reviewers can learn from the mistake.
Link to the comment on the PR: N/A
-
[Contributor] If the regression was CRITICAL (e.g. interrupts a core flow) A discussion in #expensify-open-source has been started about whether any other steps should be taken (e.g. updating the PR review checklist) in order to catch this type of bug sooner.
Link to discussion: N/A
-
[Contributor] If it was decided to create a regression test for the bug, please propose the regression test steps using the template below to ensure the same bug will not reach production again.
Regression Test Proposal
Precondition:
- User A has
focusmode and DM with user B. - Open the app as User A and User B from different devices (or from incognito and not incognito windows)
Test:
- As User A sends the message 'Test' to User B in DM, and switches to another chat
- As User B, reply in a thread to the message 'Test'
- From User A, see the reply arrives and is displayed as unread in the Inbox (All/Unreads tabs)
Do we agree 👍 or 👎
-
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsCRITICAL
Problem
Over the past few days or weeks, chats with new messages sometimes don't appear as unread in the LHN. The unread state shows only after the user opens the chat manually.
Details
Next steps
Reported in Slack.
Upwork Automation - Do Not Edit
Issue Owner
Current Issue Owner: @dmkt9