I started captioning everything about three years ago for accessibility reasons and because a large share of viewing happens with sound off. What I did not anticipate was that reading my own words on screen would change how I wrote them.

Spoken filler is invisible until it is text

The first thing you see is how much of what you say is not content.

Sort of, kind of, basically, essentially, I think, you know, obviously. In speech these pass unnoticed and they carry some conversational function. As text on screen they are a wall of noise between the viewer and the point.

I went through a year of captions and counted. Roughly one word in eight was doing nothing. On a ten minute video that is over a minute of nothing.

The caption file is a transcript of your own speech habits, and reading it is uncomfortable in a way that listening to yourself is not, because text does not have the tone that made it acceptable.

Sentence length shows up as a rhythm problem

Captions break into chunks of a few words, displayed for a couple of seconds each.

A long, subordinate-clause-heavy sentence becomes six or seven caption cards, and the viewer loses the thread of the sentence before it resolves. It reads as incoherent even though it made sense as speech.

Short declarative sentences caption beautifully. Which pushed me toward writing shorter sentences generally, and the spoken version improved too.

The general lesson is that a sentence structure which works when held in the ear does not necessarily survive being cut into pieces on screen.

The pace of speech becomes visible

Something I could not have noticed otherwise. There are passages in my videos where the captions change slowly and passages where they fly past.

The fast ones are almost always the parts where I was less sure of what I was saying and was covering by talking quickly. They are also, reliably, the parts where retention dips.

Slowing those down was one of the more effective changes I have made and I only identified them because the caption timing made it visible.

Automatic captions are a diagnostic tool

Where automatic transcription fails is informative independently of whether you use the output.

It mishears words that are unclear. It mishears technical terms said too quickly. It produces nonsense where I trailed off or where two ideas collided.

Every one of those is a place where a human listener is also working harder than they should be. The machine failing is a signal about the recording, not just about the machine.

I now run automatic captions early, before editing properly, and treat the error list as a note about which lines to re-record.

Writing for the eye and the ear at once

Practical changes that came out of it.

Numbers and technical terms get said clearly and slowly, because they caption badly and because they are what people rewind for.

Anything essential does not get said over a visual that demands attention, because the viewer cannot read captions and study a diagram simultaneously.

Captions get positioned away from anything important in the lower frame, which means composing with that region kept clear.

And lists get stated with clear numbering, because a list in captions without markers is indistinguishable from continuous prose.

The accessibility point, properly

I framed this as a craft improvement and the primary reason should be stated directly.

A substantial number of people cannot use your video without captions. Automatic captions are frequently poor enough to be genuinely unusable for someone relying on them entirely — the errors are not amusing when they are your only access to the content.

Correcting them takes about fifteen minutes for a ten minute video with a decent editor. That is a small cost for making the work usable by people who currently cannot use it.

The craft benefits are real and they are a side effect.

What I would suggest trying

Take something you made a year ago, generate a transcript, and read it as text with the video off.

It is a genuinely unpleasant experience the first time and it is the most direct feedback on your own writing available.

Most of what you will find is not errors of fact or structure. It is redundancy, filler, and sentences that never quite finished, all of which are invisible in speech and obvious on the page.

Translation makes it worse and clearer

A later stage that pushed the same lesson further. I had some work subtitled into two other languages by people who do it professionally.

The questions they came back with were all about places where my meaning was genuinely ambiguous. Idioms that do not travel. Sentences where the referent of a pronoun was unclear. Jokes that depended entirely on a phrase's second meaning.

None of those had been flagged by any English viewer, because English viewers were filling in the gaps without noticing.

A translator cannot fill in a gap. They have to ask, and every question is a place where the original was less precise than it seemed.

What I would fix first

If someone wanted to do one thing from all this, I would say: cut the filler and slow down the fast bits. Those two changes account for most of what I gained, and both are visible in a transcript within about five minutes of reading it.