Why training sticks better in your ears#

Move the words from the slide into the learner's ears and, in most corporate training, comprehension improves. This is one of the most replicated findings in instructional design, and it has a name: the modality effect.

The reason is architectural rather than preferential. Human working memory processes information through two partially separate channels, one auditory and one visual, each with its own limited capacity. When a learner reads text on screen while also studying a diagram, both tasks compete for the same visual channel and the channel saturates. Move the explanation into audio, and the words travel through the auditory channel while the visual channel handles the graphic. Nothing about the learner changed. You simply stopped forcing two jobs through one pipe.

Richard Mayer states this as the modality principle: People learn better from graphics with spoken narration than from graphics with printed on-screen text. The Cambridge Handbook of Multimedia Learning puts the consequence precisely: presenting some information visually and other information in auditory mode can expand effective working memory capacity and reduce excessive cognitive load.

That matters because of what corporate training usually looks like. The default format is a slide deck carrying blocks of text alongside a diagram or screenshot, and it is precisely the format the modality principle predicts will underperform. Against that baseline, which is the realistic comparison for most L&D teams, audio wins on established evidence.

But not because you have "auditory learners"#

Before going further, one idea needs to be cleared out of the way, because building a program on it will waste your budget.

The belief that people have a preferred sensory channel, and that matching instruction to it improves learning, is called the meshing hypothesis. It is accepted as fact by well over 90 percent of teachers in some surveys, and It has no credible scientific support.

The landmark review is Pashler, McDaniel, Rohrer, and Bjork's 2008 paper in Psychological Science in the Public Interest. They specified what valid evidence would require: classify learners by style, randomly assign different instructional modes, measure everyone identically, and find a crossover interaction where each group performs best under its matched condition. They found abundant evidence that people have preferences about presentation. They found no published studies meeting that standard.

Later work is blunter. Riener and Willingham wrote in 2010 that there is no credible evidence that learning styles exist. Paul Kirschner's 2017 paper in Computers & Education is titled "Stop propagating the learning styles myth," arguing the idea is not merely wasteful but actively harmful because it diverts budget from strategies that work.

There is a measurement problem underneath it too. Learning style inventories rely on self-report; those self-assessments are generally unvalidated, and people are poor judges of their own learning. What the instruments capture is preference, not capability.

This distinction matters commercially. If a vendor pitches you audio on the grounds that some of your people are auditory learners, they are selling a debunked premise, and you should apply that test to us as well. The real argument for audio is the cognitive load argument above, and it is stronger because it applies to everyone rather than to a third of your workforce.

The mechanism: cognitive load#

The framework underneath all of this is Cognitive Load Theory, developed by John Sweller and extended to multimedia by Mayer and Roxana Moreno.

Working memory is severely limited, and those limits are relatively inflexible. Instructional design is essentially the management of that constraint. Load comes in three forms: intrinsic load from the inherent complexity of the material, extraneous load created by how you present it, and germane load, the productive effort that actually builds understanding.

Only extraneous load is fully within your control, and it is where most training programs lose. Every avoidable demand you place on working memory is capacity taken away from the germane processing that produces learning. The modality effect works because it reduces extraneous load rather than because it makes anyone smarter.

The redundancy trap#

Here is where most organizations accidentally cancel the benefit they were trying to capture.

Mayer's redundancy principle: graphics plus narration outperforms graphics plus narration plus on-screen text. Presenting identical information simultaneously through both channels does not reinforce it. It consumes working memory that is then unavailable for learning.

The practical implication is uncomfortable, because it is close to universal practice: narrating the text on your slides makes your training worse. Not neutral, worse. If your e-learning module displays a paragraph while a voice reads that paragraph aloud, you have added extraneous load and degraded comprehension.

Audio has to replace on-screen text, not accompany it. The slide carries the visual, the audio carries the words, and the two do different jobs. That single change is the highest-return fix available to most L&D teams, and it costs nothing but a redesign.

Conversational beats formal#

A second principle with an unusually good effort-to-return ratio.

Mayer's personalization principle: people learn better from multimedia when the words are in conversational style rather than formal style. "Here is the thing most people get wrong about this" outperforms "It should be noted that a common misconception exists."

Formal register is not more professional in any way that matters to learning outcomes. It is simply harder to process, and the extra processing comes out of the same limited budget. This is one reason two-host conversational formats tend to outperform single-voice narration reading a document: the register is closer to how people actually explain things to each other.

Pre-training: audio before the lesson#

The principle most underused in corporate onboarding, and possibly the highest-leverage one for technical and industrial roles.

The pre-training principle: people learn better from a lesson when they already know the names and characteristics of the key concepts going in. Front-loading terminology and core concepts measurably improves what the main training event achieves, because learners are not simultaneously decoding vocabulary and absorbing procedure.

For roles dense in specialist language, which describes most manufacturing, healthcare, logistics, and regulated environments, a short audio primer delivered before a course or before a new hire's first day does real work. It is exactly what preboarding audio is for, and we cover that application in detail in our guide to cutting onboarding ramp time with a welcome podcast.

The conditions where the advantage stops#

Being straight about the boundaries, because the modality effect is conditional rather than universal, and knowing the conditions is what lets you apply it well.

Spoken information is transient. Written information persists. Once a sentence has been spoken it is gone unless the listener holds it in working memory. Written text stays on the page, can be re-read, and can be consulted while processing the next idea.

This produces the transient information effect, and in some conditions it reverses the modality advantage. Singh, Marcus, and Ayres demonstrated that written text produced larger learning gains than identical spoken text, attributing the difference to the extra load that longer spoken passages create precisely because they lack permanence.

So the advantage holds when segments are short, when audio accompanies a visual, and when pacing suits the learner. It weakens when spoken passages run long, when material is dense and unfamiliar, and when learners need to control their own pace and re-read. Our comparison piece on audio versus text retention covers that trade-off in more depth.

The rule that falls out is clean: audio carries meaning well and carries precision badly.

Six principles you can apply this week#

None of these require new software.

1. Stop narrating on-screen text. Pick one channel per piece of information. Use narration to explain the visual, never to duplicate written words.

2. Split channels rather than duplicating them. When you have a diagram, process flow, or screenshot, put the explanation in audio and keep the visual clean.

3. Keep spoken segments short. Because spoken content is transient, load accumulates across long passages. Break audio into units of a few minutes with clear boundaries, and let learners control pacing where possible.

4. Write scripts conversationally. Rewrite formal register into how you would explain it to a colleague. Costs nothing, works immediately.

5. Pre-train the vocabulary. Deliver terminology and core concepts before the main module or the first day, not during it.

6. Cut the interesting extras. Mayer's coherence principle: engaging but non-essential material reduces learning. The anecdote that does not serve the objective is spending working memory you need elsewhere. Scheduling matters as much as design, and the spacing intervals are in microlearning: how to train teams in small bites.

When to use audio, and when not to#

Strong candidates:

  • Explanation accompanying diagrams, processes, or demonstrations
  • Conceptual context and background, the "why this exists" layer
  • Pre-training and vocabulary priming ahead of a main module
  • Culture, norms, and tacit knowledge that resists being written down
  • Content for people who are not at a screen, including frontline and field roles
  • Anything currently going unread, where the realistic alternative is zero consumption:The clearest example of content going unread is mandatory training, covered in compliance training people actually complete.

Poor candidates:

  • Reference material people need to look up later
  • Precise figures, thresholds, tolerances, dosages, or contractual terms
  • Complex procedures where exact sequence matters
  • Anything requiring a signature or forming part of a compliance record
  • Dense unfamiliar technical content in long uninterrupted passages

If you are applying this to new hires specifically, the wider process context sits in our complete guide to employee onboarding, where the same tension appears as the compliance trap.

Where Sprep fits#

We build a document-to-podcast tool, so read this as an interested party talking rather than neutral advice.

Two of the principles above map directly onto how Sprep is built, which is worth stating openly rather than implying. Output is a two-host conversation rather than single-voice narration, which is the personalization principle in practice. And the most common use in onboarding, giving new hires context before day one, is the pre-training principle in practice.

Mechanically: Sprep converts documents you already have, PDFs, decks, and Word files, into short conversational episodes. A person reviews and approves every script before audio is generated, which matters for training content because accuracy is not optional and one-shot generation tends to introduce small errors that only surface on reading. An approved master script generated in over 70 languages on Team plans, and output distributed to Slack, Microsoft Teams, an LMS, or a private feed. DΓ€twyler uses it for onboarding and internal communications.

Where it is the wrong tool, by the same science: reference material, precise figures, and complex procedural detail belong in writing, and converting them to audio works against the transient information effect rather than with it. Audio is the comprehension and context layer alongside your written record, not a replacement for it.

FAQ#

Why does training stick better in audio? Because of the modality effect. Working memory processes auditory and visual information through partially separate channels. When words sit on screen next to a diagram, both compete for the visual channel. Moving the words to audio frees visual capacity for the graphic, which reduces cognitive load and improves comprehension. It works for everyone, not for a subset of "auditory learners."

Are some people auditory learners? People have presentation preferences, but there is no credible evidence that matching instruction to a preferred sensory channel improves outcomes. The 2008 review by Pashler and colleagues found no studies meeting the required evidential standard, and researchers including Kirschner have argued the myth actively diverts resources from effective strategies.

What is the modality principle? Mayer's modality principle states that people learn better from graphics with spoken narration than from graphics with printed on-screen text. Because the two channels process separately, moving verbal information to audio expands effective working memory capacity and reduces excessive cognitive load.

Why is narrating on-screen text bad? It violates the redundancy principle. Presenting identical information through both channels simultaneously does not reinforce learning, it consumes working memory needed for actual processing. Graphics plus narration beats graphics plus narration plus text, so audio should replace on-screen words rather than accompany them.

What is cognitive load theory? Developed by John Sweller, it holds that working memory is severely limited and instructional design should manage three load types: intrinsic load from material complexity, extraneous load from presentation design, and germane load from productive learning effort. Reducing extraneous load leaves more capacity for germane processing.

When does audio stop working for training? When spoken passages run long, when material is dense and unfamiliar, and when learners need to control pace or re-read. This is the transient information effect: speech does not persist, so longer audio accumulates cognitive load. Research by Singh, Marcus, and Ayres found written text outperformed identical spoken text for this reason.

How long should audio training segments be? Short, typically a few minutes per segment with clear boundaries. Because spoken content is transient, load builds across long uninterrupted passages. Segmentation also gives learners pacing control, which offsets audio's main structural disadvantage against text.

Should training audio be scripted formally or conversationally? Conversationally. Mayer's personalization principle finds people learn better when words are in conversational rather than formal style. This is among the cheapest available improvements, since it requires only a rewrite rather than new production or technology.

What is the pre-training principle? It states that people learn better when they already know the names and characteristics of key concepts before the main lesson. Practically, deliver a short primer covering terminology ahead of a training module or a new hire's first day. It is especially valuable in technical and industrial roles with heavy specialist vocabulary.

Can audio replace written training materials? No. Reference material, precise figures, complex procedures, and anything requiring a signature or forming part of a compliance record must stay written and searchable. Audio works best as a comprehension and context layer alongside written material, particularly for content currently going unread.

Try it on one training document#

If you want to test whether a conversational audio primer improves what your training actually achieves, the free plan converts one document into a full episode of up to 15 minutes, with the script editor included so you can review and correct every word before audio is generated.

Start free

See it in action

Convert your own documents into podcasts