---
title: "Audio vs Text: Which Drives Better Retention? (Research)"

slug: audio-vs-text-which-drives-better-retention-research

date: 2026-07-16

category: Comparison

readTime: 10 min read

excerpt: The honest answer isn't "audio is just as good as reading." It depends on what you're trying to retain, and the research is more detailed than the podcast industry usually admits.

heroImage: audio-vs-text-retention-hero.jpg

draft: true
---

## The short answer

For simple, narrative material, listening and reading produce roughly the same comprehension. For dense, technical, or unfamiliar material, reading tends to win, sometimes by a wide margin. Neither format is universally "better." The right choice depends on what you're asking people to retain, not on a blanket claim about audio versus text.

This matters beyond trivia. It's the actual evidence behind [the unread problem](https://sprep.ch/blog/posts/the-unread-problem-why-nobody-finishes-your-documents/), the idea that a format nobody finishes delivers less value than a slightly-less-perfect format people actually complete. If you're deciding whether to turn a document into a podcast, the research below tells you when that trade is worth making and when it isn't.

## Where the evidence says they're about equal

The most-cited study here is Beth Rogowsky's 2016 research at Bloomsburg University of Pennsylvania. Rogowsky split participants into three groups: one read sections of Laura Hillenbrand's nonfiction book *Unbroken* on an e-reader, one listened to the audiobook, and one did both simultaneously. All three groups took the same comprehension quiz afterward, and scores came out statistically indistinguishable across groups. Rogowsky, who described herself as a skeptic of audiobooks going in, found no comprehension penalty for listening.

That result holds up in more recent work too. A 2019 study comparing short text and audio passages found comprehension scores within two points of each other, 53 percent for text readers and 55 percent for listeners, close enough to call a tie. And in 2025 research on audiobook consumption habits, listeners retained the same amount of information as readers using a Kindle-style device, regardless of format.

So for a memoir chapter, a news story, or a straightforward narrative update, the format you choose can come down to preference and convenience. The comprehension data doesn't meaningfully favor either one.

## Where reading pulls ahead

The picture changes once the material gets denser. An earlier, widely cited study by Daniel and Woody found that when the content shifted from simple to complex, audio-only listeners scored up to 28 percent lower on comprehension than readers working through the same material. A 2025 study on reading-while-listening found a similar pattern from a different angle: conditions that included reading, whether reading alone or reading combined with audio, outperformed listening-only conditions on comprehension, with individual differences in working memory capacity explaining part of the gap.

The mechanism is not mysterious. Stephanie Del Tufo, a language scientist at the University of Delaware, explains it as a difference in how the brain has to work. Reading lets you set your own pace, reread a confusing sentence, and lean on visual structure like punctuation and paragraph breaks to organize meaning. Listening runs at the speaker's pace, and because speech is a continuous stream of blended sounds rather than neatly separated words, your working memory has to hold onto what you just heard while the next sentence keeps coming. For simple material, that's a manageable load. For a dense whitepaper or a technical report, it stacks up fast.

## The property that explains every result above

If you want one concept that predicts when audio wins and when it loses, it is **transience**.

Written text persists. It sits on the page, and the reader can look back at the previous sentence while processing the current one, scan ahead to see where the argument is going, or stop entirely and return an hour later without losing their place. Spoken information does none of that. Once a sentence has been said, it is gone unless the listener was holding it in working memory at the moment it passed.

That single difference generates the whole pattern in the research. Simple narrative material makes light demands on working memory, so transience costs little and the formats tie. Dense material makes heavy demands, so transience costs a lot and reading pulls ahead. It also explains why the gap widens as passages get longer: load accumulates across an uninterrupted stream in a way it does not across a page you can navigate.

The practical version is short. **Audio carries meaning well and carries precision badly.** Anything a person will need to look up, check, or hold alongside something else belongs in writing.

## What happens at 1.5x and 2x speed

Worth addressing directly, because a large share of podcast listening happens at increased playback speed and the effect is not neutral.

Research on accelerated speech generally finds comprehension holds up reasonably well up to around 1.5x for straightforward material, then degrades more sharply beyond it, with the decline arriving earlier for complex or unfamiliar content. The mechanism is the same one described above. Speeding up playback compresses the time available to process each sentence before the next arrives, which increases the working memory load that transience already creates.

The practical implication for anyone producing audio: assume some of your audience is listening faster than you recorded, and that the fastest listeners are absorbing the least. This is an argument for shorter segments, cleaner structure, and explicitly signposting the parts that matter, rather than an argument against audio.

## The variable neither side of this debate accounts for

Here's what rarely comes up in audio versus reading comparisons: how much of the material people actually finish. All the studies above measure comprehension among people who completed the material in a lab setting, under instruction, with no competing demands. That is not how workplace reading happens. Attention at work is already oversubscribed before your content arrives, as the numbers in [information overload at work](https://sprep.ch/blog/posts/information-overload-at-work-stats-and-solutions-for-2026/) make clear.

Outside a research lab, only 20 to 30 percent of readers make it to the end of a long document, while podcast episodes average 70 to 80 percent completion. A document that's abandoned a quarter of the way through delivers zero retention on the other three-quarters, no matter how well the person understood the part they read. A podcast finished at a slight comprehension disadvantage can still land more total information, simply because more of it reaches the listener.

That's not a rebuttal of the comprehension research above. It's because comprehension research alone is the wrong lens for a real-world decision. Format choice is really two separate questions: which format will someone finish, and does the material demand the pacing control that reading provides. Sometimes the honest answer is both.

## Working the numbers on a real document

The abstract version of that argument is easy to nod along to and easy to dismiss. Here is what it looks like with figures attached.

Take a 3,000 word policy update sent to 500 employees, and assume the comprehension research applies at its least favourable to audio, a 28 percent penalty on dense material.

**Written version.** Suppose 25 percent of recipients reach the end, which is generous for a long internal document. That is 125 people who absorbed the full content at full comprehension. Everyone else got some fraction of the opening and nothing after that.

**Audio version.** Suppose 75 percent complete it, at the low end of typical podcast completion. That is 375 people, each retaining roughly 72 percent of what a full reader would have retained, applying the worst-case penalty.

The effective totals are 125 full-comprehension equivalents for text against about 270 for audio. Audio delivers more than twice the total information received, while being measurably worse per person who finishes.

The numbers are illustrative rather than a benchmark, and you should substitute your own completion data. But the shape holds across a wide range of plausible inputs, and it only reverses when written completion is unusually high or the material is precise enough that partial comprehension is dangerous rather than merely imperfect. Which is exactly the case where you keep it written.

## Individual differences that change the answer

The averages hide real variation, and three factors matter enough to plan around.

**Working memory capacity.** The 2025 reading-while-listening study found that individual differences in working memory explained part of the performance gap. People with more available capacity handle transience better, which means the audio penalty on complex material falls harder on the people already stretched.

**Prior knowledge.** Someone familiar with the subject can predict where a sentence is going and reconstruct anything they missed. Someone encountering the terminology for the first time cannot, which makes audio a poor vehicle for genuine first exposure to unfamiliar technical material, and a good one for reinforcing something already introduced.

**Language proficiency.** Non-native speakers generally lose more from transience than native speakers, since decoding takes marginal capacity that is then unavailable for comprehension. For multilingual workforces this argues for audio in the listener's own language rather than audio in the corporate language, and for keeping a written version available alongside.

None of these change the overall pattern. They change who is affected by it, which matters when you are choosing a format for a whole organisation rather than for yourself.

## A caution about the research itself

Two limits worth knowing before you lean too hard on any single number in this article.

Most of these studies measure comprehension shortly after exposure, usually with multiple-choice questions. That measures recognition more than durable retention, and the two can diverge. The 2019 study that used free-recall questions rather than multiple choice found retention was measurably higher for text even where comprehension scores looked similar, which suggests the format gap may be wider than recognition-based testing shows.

The sample sizes and materials also vary enormously, from short passages to book chapters, across student and general adult populations. That is why the numbers quoted here differ between studies and why treating any one figure as precise would be a mistake. The direction of the findings is consistent. The magnitudes are not.

## So which should you use?

A rough guide based on the research above:

- **Simple, narrative, or familiar material** such as updates, stories, and general announcements: comprehension is a wash. Choose based on what people will actually finish.
- **Dense, technical, or unfamiliar material** such as policy detail, financial figures, and anything with numbers people need to act on precisely: reading has a real edge, or a written reference alongside an audio version so people can double-check the parts that matter.
- **Anything currently going unread**, such as the report sitting at 20 percent scroll depth: a completion-optimized format beats a comprehension-optimized one nobody opens.
- **Anything requiring a signature, or forming part of a compliance record:** written, always, with audio as an optional comprehension layer on top.

This is also why a document-to-podcast script matters more than raw text-to-speech. A well-restructured script can slow down for the genuinely complex parts, add the context a listener can't get by flipping back a page, and keep pacing close to what a reader would choose for themselves, closing some of the gap the research above documents. Sprep's scripts go through human review before audio generation for exactly this reason: getting the pacing and framing right on the dense parts is where audio-only tools tend to lose the comprehension edge.

If the underlying problem you are solving is that internal updates are not landing at all, the organisational version of this trade-off is covered in [why employees don't read company updates](https://sprep.ch/blog/posts/why-employees-dont-read-company-updates-data-and-fixes/).

## FAQ

**Does listening to a podcast help you retain information as well as reading?**
For simple or narrative content, yes, research finds comparable comprehension scores. For dense or technical content, reading tends to produce better recall, sometimes by as much as 28 percent in controlled studies. The gap tracks how much working memory the material demands rather than anything inherent to ears versus eyes.

**Is reading always better than listening for learning?**
No. The comprehension gap only shows up reliably with complex material. For straightforward content, multiple studies including Rogowsky's 2016 work find no meaningful difference between reading and listening, and format can be chosen on convenience alone.

**Why does reading work better for complex material?**
Because written text persists and speech does not. Reading lets you control pace, reread confusing passages, and use visual structure to organize meaning. Listening runs at the speaker's pace and puts more load on working memory, since each sentence is gone once it has passed.

**What is the transience effect?**
It is the property that spoken information disappears once said, while written text remains available to re-read. Transience is the single mechanism that explains why the two formats tie on simple material and diverge on complex material, since the cost of losing a sentence rises with how much you need to hold onto it.

**Does listening at 1.5x or 2x speed hurt comprehension?**
Generally yes, though the effect is modest up to around 1.5x for straightforward material and grows sharper beyond that, arriving earlier for complex content. Faster playback compresses processing time per sentence, which compounds the load transience already creates.

**If reading has better comprehension, why convert documents to audio at all?**
Because comprehension only counts for the part people actually finish. Long documents see roughly 20 to 30 percent completion, while podcast episodes average 70 to 80 percent. A format people finish can deliver more total retained information than one with better comprehension but a much lower completion rate.

**Does reading while listening at the same time improve retention?**
Rogowsky's study found the dual-modality group performed comparably to the single-format groups rather than better, so simultaneous reading and listening is not a reliable upgrade. Later work found reading-inclusive conditions outperformed listening-only, suggesting the reading component rather than the combination is what carries the benefit.

**Who is most affected by the audio comprehension penalty?**
People with less available working memory, people encountering the subject for the first time, and non-native speakers processing in a second language. The averages in the research hide meaningful variation, which matters when choosing a format for an entire organisation rather than for yourself.

**How reliable are audio versus text comprehension studies?**
Directionally consistent, numerically not. Most rely on multiple-choice testing shortly after exposure, which measures recognition rather than durable retention, and sample sizes and materials vary widely. One study using free-recall questions found a larger advantage for text than recognition-based testing suggested.

**When should content stay written no matter what?**
Anything people need to reference precisely, anything containing figures or thresholds they will check, anything requiring a signature, and anything forming part of a compliance record. Audio can sit alongside these as a comprehension layer, but it should never replace the written version of a record.

<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [
{"@type":"Question","name":"Does listening to a podcast help you retain information as well as reading?","acceptedAnswer":{"@type":"Answer","text":"For simple or narrative content, yes, research finds comparable comprehension scores. For dense or technical content, reading tends to produce better recall, sometimes by as much as 28 percent in controlled studies. The gap tracks how much working memory the material demands rather than anything inherent to ears versus eyes."}},
{"@type":"Question","name":"Is reading always better than listening for learning?","acceptedAnswer":{"@type":"Answer","text":"No. The comprehension gap only shows up reliably with complex material. For straightforward content, multiple studies including Rogowsky's 2016 work find no meaningful difference between reading and listening, and format can be chosen on convenience alone."}},
{"@type":"Question","name":"Why does reading work better for complex material?","acceptedAnswer":{"@type":"Answer","text":"Because written text persists and speech does not. Reading lets you control pace, reread confusing passages, and use visual structure to organize meaning. Listening runs at the speaker's pace and puts more load on working memory, since each sentence is gone once it has passed."}},
{"@type":"Question","name":"What is the transience effect?","acceptedAnswer":{"@type":"Answer","text":"It is the property that spoken information disappears once said, while written text remains available to re-read. Transience is the single mechanism that explains why the two formats tie on simple material and diverge on complex material, since the cost of losing a sentence rises with how much you needed to hold onto it."}},
{"@type":"Question","name":"Does listening at 1.5x or 2x speed hurt comprehension?","acceptedAnswer":{"@type":"Answer","text":"Generally yes, though the effect is modest up to around 1.5x for straightforward material and grows sharper beyond that, arriving earlier for complex content. Faster playback compresses processing time per sentence, which compounds the load transience already creates."}},
{"@type":"Question","name":"If reading has better comprehension, why convert documents to audio at all?","acceptedAnswer":{"@type":"Answer","text":"Because comprehension only counts for the part people actually finish. Long documents see roughly 20 to 30 percent completion, while podcast episodes average 70 to 80 percent. A format people finish can deliver more total retained information than one with better comprehension but a much lower completion rate."}},
{"@type":"Question","name":"Does reading while listening at the same time improve retention?","acceptedAnswer":{"@type":"Answer","text":"Rogowsky's study found the dual-modality group performed comparably to the single-format groups rather than better, so simultaneous reading and listening is not a reliable upgrade. Later work found reading-inclusive conditions outperformed listening-only, suggesting the reading component rather than the combination is what carries the benefit."}},
{"@type":"Question","name":"Who is most affected by the audio comprehension penalty?","acceptedAnswer":{"@type":"Answer","text":"People with less available working memory, people encountering the subject for the first time, and non-native speakers processing in a second language. The averages in the research hide meaningful variation, which matters when choosing a format for an entire organisation rather than for yourself."}},
{"@type":"Question","name":"How reliable are audio versus text comprehension studies?","acceptedAnswer":{"@type":"Answer","text":"Directionally consistent, numerically not. Most rely on multiple-choice testing shortly after exposure, which measures recognition rather than durable retention, and sample sizes and materials vary widely. One study using free-recall questions found a larger advantage for text than recognition-based testing suggested."}},
{"@type":"Question","name":"When should content stay written no matter what?","acceptedAnswer":{"@type":"Answer","text":"Anything people need to reference precisely, anything containing figures or thresholds they will check, anything requiring a signature, and anything forming part of a compliance record. Audio can sit alongside these as a comprehension layer, but it should never replace the written version of record."}}
]
}
</script>

## Get more research like this

This post is part of our ongoing series on why workplace content goes unread and what the evidence actually supports. Subscribe to get the next one, delivered as both a short read and a Sprep-generated audio version.

[Subscribe to the newsletter](/newsletter)

See it in action

Convert your own documents into podcasts