Every real-time thing I have built — the multiplayer prototypes, the live odds feed on a betting app, a ride-hailing driver map — eventually hit the same bug: the socket drops, the client reconnects badly, and either the server gets hammered or the user sits on stale data without knowing it. The naive fix is ws.onclose = () => connect(). This post is what I replace it with: a small state machine that knows which closes to retry, backs off with full jitter, notices dead connections with a heartbeat, and resumes state instead of pretending nothing was missed.
Why the naive reconnect breaks
Reconnecting immediately inside onclose has three failure modes that only show up in production:
- Hot loop. If the server rejects the handshake (bad token, server down),
onclosefires immediately, you reconnect, it fails again — hundreds of attempts a second per tab. - Thundering herd. When a server restarts, every client disconnects at the same instant. With a fixed delay (or even plain exponential backoff) they all come back in synchronised waves, and the reconnect spike takes the fresh server straight back down.
- Silent gaps. Messages sent while you were disconnected are simply gone. The UI looks live but is missing events.
And there is a fourth that does not even fire onclose: the half-open connection. A phone switches from Wi-Fi to mobile data, a NAT drops the mapping, and the TCP socket stays “open” on your side for minutes while nothing arrives. readyState still says OPEN.
Step 1: classify the close code
Not every close deserves a retry. The CloseEvent.code tells you most of what you need (RFC 6455 defines the 1000-range; 4000–4999 is yours to use):
- 1000 Normal closure — someone meant to close. If it was you (logout, unmount), stop. If the server sends it, treat it as “go away” and stop unless your protocol says otherwise.
- 1001 Going away — server shutting down or redeploying. Retry, with backoff.
- 1006 Abnormal closure — never sent on the wire; the browser reports it when the TCP connection died without a close frame. It is also what you get when the HTTP upgrade itself fails (401, 502), because browsers deliberately hide handshake details. Retry, with backoff.
- 1008 Policy violation and 1003 — the server rejected what you sent. Retrying the same thing will fail the same way. Stop.
- 1011 Internal error, 1012 Service restart, 1013 Try again later — retry, with backoff (for 1013, a longer floor).
- 4000–4999 — define your own. I use
4401for “token expired”: the client refreshes auth, then reconnects immediately, once. Without that, an expired token turns into an infinite 1006 loop, because the browser cannot see the 401.
That last point is the single most useful thing in this post: do auth failures after the upgrade, not during it. Accept the socket, validate the token in the first message, and close with a 4xxx code the client can read.
Step 2: full-jitter exponential backoff
Plain exponential backoff (1 s, 2 s, 4 s…) still synchronises clients, because they all started failing at the same moment. The fix, from AWS’s well-known backoff analysis, is full jitter: pick a random delay between zero and the exponential cap.
delay = random(0, min(cap, base × 2^attempt))
With base = 500 ms and cap = 30 s, attempt 0 waits 0–0.5 s, attempt 3 waits 0–4 s, attempt 6 and beyond wait 0–30 s. Full jitter spreads the herd across the whole window, which minimises the peak reconnect rate — the number that actually overloads an accept queue. “Equal jitter” (half fixed, half random) keeps a minimum wait but produces a higher peak; I only use it when I want a guaranteed floor.
The subtle bug almost every hand-rolled version has: resetting the attempt counter in onopen. If the server accepts the connection and then immediately drops it (crash loop, overloaded, auth-in-first-message fails), you reset to attempt 0 every time and are back to a hot loop. Reset only after the connection has stayed open for a while — I use 10 seconds — or after the first real application message arrives.
Step 3: heartbeat to catch half-open sockets
The browser WebSocket API does not expose protocol-level ping/pong frames. The browser answers server pings automatically, which keeps the server honest, but your client code cannot send a ping or see a pong. So for client-side detection you need an application-level heartbeat:
- Server sends
{"t":"ping"}every 20–25 s (under the typical 30–60 s idle timeouts of load balancers and proxies). - Client resets a watchdog timer on any incoming message. If nothing arrives for ~2.5× the interval, call
ws.close()and move to WAITING yourself. - Also listen for
onlineandvisibilitychange: when the tab comes back or the network returns, check the watchdog immediately instead of waiting out a 30 s backoff.
Note that calling close() on a half-open socket can take a while to fire onclose. Do not wait for it: detach the handlers, drop the reference and start the reconnect timer straight away.
Step 4: resume state, don’t just reconnect
A reconnected socket is a new session. Anything published in between is lost unless your protocol handles it. The simplest scheme that works:
- The server stamps every message with a monotonically increasing
seqper stream. - The client stores the last
seqit processed. - On reconnect, the first message is
{"t":"resume","token":…,"lastSeq":1842}. - The server replays 1843 onward from a short ring buffer — or, if the gap is older than the buffer, replies
{"t":"snapshot"}with full current state. - The client drops any message with
seq <= lastSeq, so duplicates during the handover are harmless.
A gap in seq during normal operation is also a signal: request a snapshot rather than rendering an inconsistent view. For odds, prices or game state, the snapshot path matters more than the replay path — it is always correct, just heavier. Outgoing messages need the same care: queue what the user did while offline, but give each one an idempotency key so a message that was sent-but-unacknowledged before the drop is not applied twice (the same idea as payment idempotency).
The code
About 60 lines, no dependency:
const RETRY = new Set([1001, 1006, 1011, 1012, 1013]);
export function reconnectingSocket(url, { onMessage, getToken, base = 500, cap = 30000, hb = 25000 }) {
let ws, attempt = 0, lastSeq = 0, stopped = false, retryT, stableT, dogT;
const delay = () => Math.random() * Math.min(cap, base * 2 ** attempt);
const dog = () => { clearTimeout(dogT); dogT = setTimeout(() => drop(1006), hb * 2.5); };
function connect() {
ws = new WebSocket(url);
ws.onopen = () => {
ws.send(JSON.stringify({ t: 'resume', token: getToken(), lastSeq }));
stableT = setTimeout(() => (attempt = 0), 10000); dog();
};
ws.onmessage = (e) => {
dog(); const m = JSON.parse(e.data);
if (m.t === 'ping') return;
if (m.seq && m.seq <= lastSeq) return; // duplicate
if (m.seq) lastSeq = m.seq;
onMessage(m);
};
ws.onclose = (e) => drop(e.code);
}
function drop(code) {
clearTimeout(stableT); clearTimeout(dogT);
if (ws) { ws.onclose = ws.onmessage = ws.onopen = null; try { ws.close(); } catch {} }
if (stopped) return;
if (code === 4401) { attempt = 0; return void setTimeout(connect, 0); } // token refreshed via getToken()
if (!RETRY.has(code)) return; // 1000, 1003, 1008 … → STOPPED
retryT = setTimeout(connect, delay()); attempt++;
}
const kick = () => { if (!stopped && ws?.readyState !== 1) { clearTimeout(retryT); attempt = 0; connect(); } };
addEventListener('online', kick);
connect();
return { send: (m) => ws?.readyState === 1 && ws.send(JSON.stringify(m)),
close: () => { stopped = true; clearTimeout(retryT); removeEventListener('online', kick); drop(1000); } };
}
In real use getToken() should return a fresh token (refresh it before calling reconnect on 4401), and 4401 needs a guard so a server that always returns it cannot create a loop — cap it at one immediate retry, then fall back to normal backoff.
Server side: don’t make the herd worse
- Drain on deploy. Close connections with 1012 in batches over 30–60 s rather than all at once; clients’ jitter only helps if the disconnects are not all simultaneous too.
- Rate-limit accepts. Answer overload with 1013 and let the client’s backoff do its job.
- Keep the replay buffer bounded by time, not count, and document it: “resume works for gaps under 2 minutes; beyond that you get a snapshot.”
Checklist
- Close codes classified: retry 1001/1006/1011–1013, stop on 1000/1003/1008.
- Auth failures sent as a 4xxx close after the upgrade, not a 401 during it.
- Full jitter, 500 ms base, 30 s cap.
- Attempt counter reset after ~10 s stable, never in
onopen. - App-level heartbeat + watchdog at 2.5× interval;
onlinetriggers an immediate check. - Sequence numbers, resume with
lastSeq, snapshot fallback, duplicate drop. - Outgoing messages carry idempotency keys.
- Show the user the state (“Reconnecting…”) — silent staleness is worse than a visible banner.