Instructional videos produce a distinctive retention shape, with abrupt dips and replays clustered at the same few seconds. Those points mark where the explanation outran the viewer.
A tutorial is watched with hands busy
Most instructional viewing happens alongside the task. The viewer is watching, then acting, then returning, which means the video is being consumed in fragments rather than continuously.
Every fragment boundary is a place where playback is interrupted deliberately. The graph is therefore not measuring interest; it is measuring how often the video asked the viewer to do something.
That reframing matters, because a dip in a tutorial can indicate the video is working. The problem is only the dips that come with a rewind rather than a pause.
Step density is the usual cause
Rewinds cluster where several actions were compressed into one sentence. A presenter who is fluent in the task performs three steps while narrating one, and the viewer loses the thread between them.
The effect is invisible to the person recording. Familiarity removes the perception of separate steps, so the section feels like a single unremarkable moment on the way to something harder.
Locating these points does not require guesswork. The rewind clusters identify them directly, and they usually sit before the section the creator expected to be difficult.
Naming things before using them
A second reliable cause is vocabulary. Using a term, a menu name or an abbreviation before defining it forces the viewer to stop and resolve it elsewhere.
Instructional writing works better when every noun is introduced once at the moment it first becomes relevant, rather than in a glossary section the viewer skipped.
On-screen text carries this load well, because it persists while the presenter continues and can be read at whatever speed the viewer needs.
Pace is a production decision
Tightening the edit removes dead air and also removes the gaps where a viewer catches up. A tutorial edited for pace can be harder to follow than the unedited take it came from.
Holding a shot for a beat after an action completes gives the viewer somewhere to land. It costs a second and prevents a rewind that costs considerably more attention.
What to do with the graph
The productive response is to re-record the identified passage rather than the whole video, breaking one dense sentence into three and letting the visual show each action separately.
Repeated dips at a chapter boundary mean something different. Those indicate viewers arriving from search, skipping to what they came for, and they call for better chaptering rather than better explanation.