WebSocket reconnect done right: backoff, jitter and resuming state

Every real-time thing I have built — the multiplayer prototypes, the live odds feed on a betting app, a ride-hailing driver map — eventually hit the same bug: the socket drops, the client reconnects badly, and either the server gets hammered or the user sits on stale data without knowing it. The naive fix is ws.onclose = () => connect(). This post is what I replace it with: a small state machine that knows which closes to retry, backs off with full jitter, notices dead connections with a heartbeat, and resumes state instead of pretending nothing was missed.

Why the naive reconnect breaks

Reconnecting immediately inside onclose has three failure modes that only show up in production:

  • Hot loop. If the server rejects the handshake (bad token, server down), onclose fires immediately, you reconnect, it fails again — hundreds of attempts a second per tab.
  • Thundering herd. When a server restarts, every client disconnects at the same instant. With a fixed delay (or even plain exponential backoff) they all come back in synchronised waves, and the reconnect spike takes the fresh server straight back down.
  • Silent gaps. Messages sent while you were disconnected are simply gone. The UI looks live but is missing events.

And there is a fourth that does not even fire onclose: the half-open connection. A phone switches from Wi-Fi to mobile data, a NAT drops the mapping, and the TCP socket stays “open” on your side for minutes while nothing arrives. readyState still says OPEN.

CONNECTING OPEN WAITING full-jitter delay STOPPED onopen 1006/1001/1011, heartbeat miss timer fires 1000 / 1008 / 4401 reset attempt counter only after OPEN is stable (~10 s)
Four states. The only interesting decisions are which closes go to WAITING vs STOPPED, and when to reset the backoff.

Step 1: classify the close code

Not every close deserves a retry. The CloseEvent.code tells you most of what you need (RFC 6455 defines the 1000-range; 4000–4999 is yours to use):

  • 1000 Normal closure — someone meant to close. If it was you (logout, unmount), stop. If the server sends it, treat it as “go away” and stop unless your protocol says otherwise.
  • 1001 Going away — server shutting down or redeploying. Retry, with backoff.
  • 1006 Abnormal closure — never sent on the wire; the browser reports it when the TCP connection died without a close frame. It is also what you get when the HTTP upgrade itself fails (401, 502), because browsers deliberately hide handshake details. Retry, with backoff.
  • 1008 Policy violation and 1003 — the server rejected what you sent. Retrying the same thing will fail the same way. Stop.
  • 1011 Internal error, 1012 Service restart, 1013 Try again later — retry, with backoff (for 1013, a longer floor).
  • 4000–4999 — define your own. I use 4401 for “token expired”: the client refreshes auth, then reconnects immediately, once. Without that, an expired token turns into an infinite 1006 loop, because the browser cannot see the 401.

That last point is the single most useful thing in this post: do auth failures after the upgrade, not during it. Accept the socket, validate the token in the first message, and close with a 4xxx code the client can read.

Step 2: full-jitter exponential backoff

Plain exponential backoff (1 s, 2 s, 4 s…) still synchronises clients, because they all started failing at the same moment. The fix, from AWS’s well-known backoff analysis, is full jitter: pick a random delay between zero and the exponential cap.

delay = random(0, min(cap, base × 2^attempt))

With base = 500 ms and cap = 30 s, attempt 0 waits 0–0.5 s, attempt 3 waits 0–4 s, attempt 6 and beyond wait 0–30 s. Full jitter spreads the herd across the whole window, which minimises the peak reconnect rate — the number that actually overloads an accept queue. “Equal jitter” (half fixed, half random) keeps a minimum wait but produces a higher peak; I only use it when I want a guaranteed floor.

The subtle bug almost every hand-rolled version has: resetting the attempt counter in onopen. If the server accepts the connection and then immediately drops it (crash loop, overloaded, auth-in-first-message fails), you reset to attempt 0 every time and are back to a hot loop. Reset only after the connection has stayed open for a while — I use 10 seconds — or after the first real application message arrives.

Step 3: heartbeat to catch half-open sockets

The browser WebSocket API does not expose protocol-level ping/pong frames. The browser answers server pings automatically, which keeps the server honest, but your client code cannot send a ping or see a pong. So for client-side detection you need an application-level heartbeat:

  • Server sends {"t":"ping"} every 20–25 s (under the typical 30–60 s idle timeouts of load balancers and proxies).
  • Client resets a watchdog timer on any incoming message. If nothing arrives for ~2.5× the interval, call ws.close() and move to WAITING yourself.
  • Also listen for online and visibilitychange: when the tab comes back or the network returns, check the watchdog immediately instead of waiting out a 30 s backoff.

Note that calling close() on a half-open socket can take a while to fire onclose. Do not wait for it: detach the handlers, drop the reference and start the reconnect timer straight away.

Step 4: resume state, don’t just reconnect

A reconnected socket is a new session. Anything published in between is lost unless your protocol handles it. The simplest scheme that works:

  1. The server stamps every message with a monotonically increasing seq per stream.
  2. The client stores the last seq it processed.
  3. On reconnect, the first message is {"t":"resume","token":…,"lastSeq":1842}.
  4. The server replays 1843 onward from a short ring buffer — or, if the gap is older than the buffer, replies {"t":"snapshot"} with full current state.
  5. The client drops any message with seq <= lastSeq, so duplicates during the handover are harmless.

A gap in seq during normal operation is also a signal: request a snapshot rather than rendering an inconsistent view. For odds, prices or game state, the snapshot path matters more than the replay path — it is always correct, just heavier. Outgoing messages need the same care: queue what the user did while offline, but give each one an idempotency key so a message that was sent-but-unacknowledged before the drop is not applied twice (the same idea as payment idempotency).

The code

About 60 lines, no dependency:

const RETRY = new Set([1001, 1006, 1011, 1012, 1013]);
export function reconnectingSocket(url, { onMessage, getToken, base = 500, cap = 30000, hb = 25000 }) {
  let ws, attempt = 0, lastSeq = 0, stopped = false, retryT, stableT, dogT;
  const delay = () => Math.random() * Math.min(cap, base * 2 ** attempt);
  const dog = () => { clearTimeout(dogT); dogT = setTimeout(() => drop(1006), hb * 2.5); };
  function connect() {
    ws = new WebSocket(url);
    ws.onopen = () => {
      ws.send(JSON.stringify({ t: 'resume', token: getToken(), lastSeq }));
      stableT = setTimeout(() => (attempt = 0), 10000); dog();
    };
    ws.onmessage = (e) => {
      dog(); const m = JSON.parse(e.data);
      if (m.t === 'ping') return;
      if (m.seq && m.seq <= lastSeq) return; // duplicate
      if (m.seq) lastSeq = m.seq;
      onMessage(m);
    };
    ws.onclose = (e) => drop(e.code);
  }
  function drop(code) {
    clearTimeout(stableT); clearTimeout(dogT);
    if (ws) { ws.onclose = ws.onmessage = ws.onopen = null; try { ws.close(); } catch {} }
    if (stopped) return;
    if (code === 4401) { attempt = 0; return void setTimeout(connect, 0); } // token refreshed via getToken()
    if (!RETRY.has(code)) return; // 1000, 1003, 1008 … → STOPPED
    retryT = setTimeout(connect, delay()); attempt++;
  }
  const kick = () => { if (!stopped && ws?.readyState !== 1) { clearTimeout(retryT); attempt = 0; connect(); } };
  addEventListener('online', kick);
  connect();
  return { send: (m) => ws?.readyState === 1 && ws.send(JSON.stringify(m)),
    close: () => { stopped = true; clearTimeout(retryT); removeEventListener('online', kick); drop(1000); } };
}

In real use getToken() should return a fresh token (refresh it before calling reconnect on 4401), and 4401 needs a guard so a server that always returns it cannot create a loop — cap it at one immediate retry, then fall back to normal backoff.

Server side: don’t make the herd worse

  • Drain on deploy. Close connections with 1012 in batches over 30–60 s rather than all at once; clients’ jitter only helps if the disconnects are not all simultaneous too.
  • Rate-limit accepts. Answer overload with 1013 and let the client’s backoff do its job.
  • Keep the replay buffer bounded by time, not count, and document it: “resume works for gaps under 2 minutes; beyond that you get a snapshot.”

Checklist

  1. Close codes classified: retry 1001/1006/1011–1013, stop on 1000/1003/1008.
  2. Auth failures sent as a 4xxx close after the upgrade, not a 401 during it.
  3. Full jitter, 500 ms base, 30 s cap.
  4. Attempt counter reset after ~10 s stable, never in onopen.
  5. App-level heartbeat + watchdog at 2.5× interval; online triggers an immediate check.
  6. Sequence numbers, resume with lastSeq, snapshot fallback, duplicate drop.
  7. Outgoing messages carry idempotency keys.
  8. Show the user the state (“Reconnecting…”) — silent staleness is worse than a visible banner.
→ WebSocket vs SSE vs WebTransport → Server-authoritative multiplayer on a budget → WebRTC DataChannel tuning → Portfolio & projects
How do I automatically reconnect a WebSocket in JavaScript?

Listen for the close event, decide from the close code whether to retry (1001, 1006, 1011-1013) or stop (1000, 1003, 1008), then schedule a new WebSocket after a full-jitter exponential backoff delay. Reset the attempt counter only after the connection has been stable for about 10 seconds.

What is full jitter backoff?

Full jitter picks a random delay between zero and min(cap, base x 2^attempt). It spreads reconnecting clients across the whole window, which minimises the peak reconnect rate after a server restart compared with plain or equal-jitter backoff.

Why do I get WebSocket close code 1006?

1006 is never sent over the wire. The browser reports it when the connection ended without a close frame, for example a network drop, a proxy timeout, or a failed HTTP upgrade such as a 401 or 502, whose details browsers hide.

Can the browser WebSocket API send ping frames?

No. Browsers answer server pings automatically but do not expose ping or pong to JavaScript. To detect half-open connections on the client, use an application-level heartbeat message and a watchdog timer.

How do I avoid missing messages after a reconnect?

Have the server number messages with a per-stream sequence, send the last processed sequence in a resume message on reconnect, replay newer messages from a bounded buffer or send a full snapshot if the gap is too old, and drop duplicates on the client.