Replay is the experiment before the experiment
Why a candidate decoder should meet yesterday's run before it meets a person, a robot, or a new piece of hardware.
A checkpoint is not evidence
A validation number tells you something useful, but it does not tell you how a model behaves inside a particular loop. The boundary cases are usually in the run itself: dropped samples, an awkward posture, a stale command, a subject doing something unexpected.
A replayable log turns those moments into a test set that still has the shape of the real system. Run the candidate against the same windows, compare its commands to the recorded baseline, and look at the cases where it changes its mind.
Keep the context with the result
The replay report should name the run, decoder versions, configuration and timing budget that produced it. Otherwise a good result becomes an anecdote as soon as the next checkpoint appears.
This is not an argument for never trying anything new. It is a way to make the first live test smaller, clearer and easier to stop.