Replay systems from an input log

The nice thing about a deterministic simulation is that you get replays almost for free. If the same inputs always produce the same state, you don't need to record frames or stream state — you record the seed and the button presses, and playback is just running the game again. A ten-minute match becomes a few kilobytes instead of a few hundred megabytes of video. The catch is the word "almost": the replay only works if nothing outside the input log can influence the simulation, and finding the one place where something does is most of the work.

What a replay actually is

A session is fully described by three things: the initial world, the RNG seed, and the input log. Everything else — positions, scores, particle timings, who won — is a pure function of those. So the recorder stores the triple and nothing else, and the player re-simulates.

That means the file format is boring, which is exactly what you want:

{ version: "sim-4.2.0", seed: "0xC0FFEE", tickRate: 60, world: "arena-3",
  players: [...], inputs: [ [tick, playerId, bitfield], ... ],
  keyframes: [ [tick, snapshotBytes], ... ], checksums: [ [tick, hash], ... ] }

The version field is the one people leave out and regret. A replay is only valid against the simulation code that produced it. Change a friction constant and every old replay quietly plays out differently — which looks exactly like a bug.

RECORD seed + world tick inputs input log ~20 KB / 10 min keyframe every 600 ticks PLAY same simulation code, fixed 60 Hz step no wall clock · no Math.random · no renderer reads state hash per tick match recorded hash? → first bad tick
Recording stores inputs, not frames. Playback re-runs the simulation and checks its own work.

Recording: bitfields and deltas

Don't record input events as they arrive from the DOM. Record the sampled input state at the start of each simulation tick — the same value the simulation consumed. Keydown timing between ticks is irrelevant, and capturing it invites drift.

Pack it small. Eight or sixteen buttons fit in one Uint16; an aim angle quantised to 1/4096 of a turn fits in another. Then delta-encode: only write a record when a player's packed input differs from their previous tick. Humans hold a key for tens of ticks at a time, so most ticks produce nothing at all. That is where the "a match fits in a few KB" number comes from.

Playback needs the inverse: for tick t, use the most recent recorded input at or before t. Store the log sorted by tick and walk it with a cursor, not a per-tick search.

The fixed timestep is non-negotiable

Every replay bug I've seen traced back to variable time entering the simulation. If your update takes dt from requestAnimationFrame, the simulation depends on the machine's frame pacing and no two playbacks agree. The simulation advances in fixed ticks — 1/60 of a second, as an exact constant — and the renderer interpolates between the last two states for display. Rendering can be as jittery as it likes; it just isn't allowed to feed anything back.

The same rule kills the whole class of "works in dev, desyncs in the tab that was backgrounded" bugs, because a background tab throttles frames but the tick loop still consumes the same number of ticks per simulated second.

Detecting desync: hash every tick

Add a cheap checksum of the simulation state at the end of every tick — a rolling FNV-1a or CRC32 over the position/velocity/health arrays is plenty. Record those hashes with the replay. On playback, compare. The first tick where they differ is where your bug is, and the difference between "the replay is broken" and "the replay is broken at tick 1,842" is an afternoon.

This is the same technique used to catch multiplayer divergence: studio postmortems from Relic's RTS titles describe CRCing state each tick across clients and dumping a circular buffer of more detailed state when a mismatch fires, because the true cause usually sits a few ticks earlier than the detection. Rollback frameworks build it in as a "sync test" mode that re-simulates the last N frames every update and compares checksums against the first run — a single-player way of finding desyncs before anyone plays online.

The recurring causes are the same list every time:

  • Wall-clock time. Date.now(), performance.now(), or a real delta anywhere inside a step.
  • Unseeded randomness. Math.random() is not reproducible and exposes no seed — use a seeded PRNG and give each subsystem its own stream.
  • Iteration order. Anything that iterates a hash-ordered collection, or sorts with a comparator that returns 0 for distinct entities. Array.prototype.sort is stable in modern engines, but only if your key is total.
  • Renderer bleed. Reading camera state, hover state, canvas size or DOM measurements inside the simulation.
  • Float transcendentals. + − × ÷ and Math.sqrt are pinned to exact IEEE 754 rounding by the spec, so they're safe. Math.sin, cos, exp and pow are only implementation-approximated — replace them with a lookup table if replays must survive across engines.

Seeking without waiting

A replay you can't scrub is half a feature. Keyframes solve it: serialise the whole simulation state every K ticks and store the snapshots alongside the log. Seeking to tick N means loading the nearest earlier keyframe and re-simulating forward with rendering disabled.

At 60 Hz, K = 600 (ten seconds) is a good default: worst case you fast-forward 600 ticks, which for a modest simulation is a few milliseconds of work and imperceptible as a scrub. Bigger worlds want smaller K; the trade is file size, and snapshots compress well because most of the state barely moves between them.

Fast-forward is also how you implement 2× and 4× playback, and how you render a replay to video headlessly: run the tick loop as fast as the CPU allows and only draw the frames you need.

Versioning, so old replays don't rot

Bump a simulation version whenever you change anything the simulation reads: constants, tuning tables, the order systems run in, the PRNG. Write it into every replay, and refuse to play a file whose version doesn't match — a clear "recorded on an older version" message beats a replay that silently diverges and makes your game look broken.

If you need long-lived replays (esports archives, dispute resolution, audit trails), keep the old simulation builds around and load the matching one, rather than trying to make one build bit-compatible with its own history. That is also the honest answer for anything money-adjacent: a replay is only evidence if the code that verifies it is the code that produced it.

What you get besides replays

Once inputs are the only source of truth, several other features fall out of the same machinery. Deterministic bug reports: a player attaches a 20 KB file and you reproduce the crash exactly. Automated regression tests: keep a folder of replays and assert the final state hash in CI, and any accidental change to physics or tuning fails the build. Spectating and killcams: the same re-simulation, just started from a keyframe a few seconds back. And the netcode path — lockstep and rollback both already transmit inputs, so the replay is a byproduct of the wire format you built anyway.

The honest cost is discipline. Determinism is not something you add at the end; it's a boundary you defend on every commit, and the per-tick checksum is what defends it for you.

Related
→ Lockstep vs rollback netcode → Fixed-point maths in JavaScript → Seeded PRNGs in JavaScript → Portfolio & projects
How does an input-log replay system work?

You record the seed, the build version and the per-tick inputs — never frames or state. Playback re-runs the same simulation over the same inputs and reproduces every frame exactly. A match is described by initial world + seed + input log, usually a few kilobytes.

Why does my replay desync partway through?

Something outside the input log leaked in: wall-clock time or a variable dt, Math.random, unordered iteration, renderer/camera reads inside a step, or float transcendentals like Math.sin. Hash the state every tick, compare against the recorded hashes, and fix the first divergent tick.

How do you seek or scrub inside a replay?

Store full state keyframes every few hundred ticks. To reach tick N, load the nearest earlier keyframe and re-simulate forward with rendering off. At 60 Hz, a keyframe every 600 ticks keeps worst-case seeks in the tens of milliseconds.

How big is an input-log replay file?

Pack each player's tick input into a bitfield and only write ticks where it changed. Because human input changes rarely relative to 60 Hz, a ten-minute two-player match usually lands in the low tens of kilobytes before compression.