Live broadcasts fail on audio far more often than on picture, and audiences abandon them faster for it. The asymmetry comes from how sound behaves in a room and in a signal chain.
Sound has no equivalent of a bad frame
A momentary picture problem passes and the viewer reconstructs what was missed, because vision tolerates interruption and the surrounding frames supply the context.
Audio interruptions destroy information outright. A dropped word cannot be inferred, and a listener who misses part of a sentence must reconstruct meaning from what follows.
This is why audiences will tolerate a soft or poorly lit picture with good sound, and will leave a sharp picture with unclear audio within a very short time.
Gain staging must be right in advance
Level set too low is buried in noise, and level set too high clips irrecoverably. Neither can be fixed once transmitted, and both are decided before anyone speaks.
The difficulty is that speech level during a rehearsal rarely matches speech level during the event, since people speak more loudly and with more variation once an audience exists.
Operators therefore leave deliberate headroom and use compression to catch peaks, accepting a slightly less dynamic sound in exchange for one that cannot destroy itself.
The room feeds back into itself
Any space with both microphones and speakers contains a loop, and if the gain around that loop exceeds a threshold the system oscillates audibly.
This is a physical property of the room rather than a fault in the equipment, which is why it depends on microphone placement, speaker position and how many microphones are open at once.
Video has no comparable phenomenon, and the fact that a whole class of failure exists only for audio explains much of the difference in difficulty.
Multiple sources must be balanced continuously
A broadcast with several speakers, music and a remote contribution requires those sources to be balanced against one another moment by moment as the event changes.
Unlike a video switch, which selects one source, an audio mix combines all of them simultaneously, so every open channel contributes noise and room sound whether or not it is being used.
Monitoring is its own problem
An operator cannot judge the broadcast mix from the room, because the room contains the live sound directly, so monitoring must be done on headphones from the output feed.
Remote contributions complicate this further, since anyone hearing their own voice returned with delay finds speaking difficult, which is why mix-minus arrangements exist at all.