Blog

SEO Gets You Found. Audio and Video Win Attention.

SEO can help publishers get discovered, but winning attention requires more. Here’s how audio, video, AI and content recycling can help every story reach a wider audience.

MediaThrive Editorial Team15 min read
Cover Image for SEO Gets You Found. Audio and Video Win Attention.
Share:LinkedInPost

Search has always been one of the foundations of digital publishing. Publishers invest heavily in SEO because discoverability matters, and that is unlikely to change. But being discovered is only the first step. A publisher can rank first in Google, win the click and still lose the reader within seconds.

That is why the conversation around publishing increasingly needs to move beyond traffic and toward attention. The question is no longer only how audiences find journalism, but also whether the format fits the way they actually consume information throughout the day.

People have not stopped wanting news, analysis or useful information. What has changed is the environment in which they consume it. Audiences move constantly between work, social platforms, messaging apps, commuting, exercise, entertainment and family responsibilities. Focused reading now competes with dozens of other activities for the same limited amount of attention.

For publishers, this creates both a problem and an opportunity. Text remains essential for search, authority, detail and reference, but it is no longer the only format audiences expect. Audio allows journalism to accompany people when looking at a screen is inconvenient or impossible, while video allows publishers to compete for attention in the feeds and platforms where audiences increasingly spend their time.

The future is therefore not about choosing between text, audio and video. It is about understanding what each format does best and allowing the same journalism to reach audiences in different contexts.

Content consumption has changed

It is easy to summarise the current shift by saying that people are reading less, but the reality is more nuanced. Readers are not necessarily less interested in information. They simply have more ways to consume it and more demands competing for their time.

A person may listen to political news while commuting, follow an entrepreneurship podcast while running, watch a short explainer during a lunch break or listen to a narrated article while cooking dinner. Later that evening, the same person may watch a longer YouTube discussion about a subject they discovered through a short video earlier in the day.

These are different consumption behaviours, but they reflect the same underlying change: audiences increasingly expect information to fit around their lives rather than requiring them to stop everything else and read.

Online audio and podcast listening have become deeply integrated into everyday routines, making listening a natural part of how people consume information throughout the day.

The numbers support that shift. According to Edison Research's The Infinite Dial 2026, 81% of Americans aged 12+ listened to online audio in the previous month and 76% had listened in the previous week. Podcast consumption also reached record levels, with 58% consuming a podcast in the previous month and 45% in the previous week.

Audio is particularly powerful because it does not require visual attention. If journalism exists only as text, it can compete only for moments when someone is willing and able to look at a screen. Adding audio extends the life of the same story into commuting, exercising, cooking, household tasks and other parts of the day in which reading is impractical.

Video addresses a different part of the same problem. People may not be able to watch while driving or running, but enormous amounts of their available screen attention have moved toward video-led environments.

The Reuters Institute's Digital News Report 2026 shows that growth in online news video is increasingly happening on social and video platforms rather than publishers' own websites. Between 2023 and 2026, weekly news-video consumption on TikTok increased by 7 percentage points and Instagram by 5 points, while consumption on publishers' websites declined by 5 points. YouTube also continues to support both short and long-form news consumption, including substantial demand for videos well beyond the typical short-form format.

That matters because video does something different from both text and audio. It can communicate an idea rapidly, make complex subjects easier to understand and create an immediate visual and emotional connection with the audience.

We have written previously about this broader multi-format behaviour in From Text to Experience: Making News Accessible in Multiple Formats. Different formats serve different consumption moments: text works well for search and depth, audio fits hands-free situations, while video is particularly effective on visual and social platforms.

The opportunity for publishers is not to replace one with another. It is to make the same journalism available in more of the moments when audiences are actually ready to consume it.

Winning attention is becoming harder

The growing influence of social and video platforms also changes how audiences first encounter journalism. People increasingly discover stories on TikTok, Instagram, YouTube, Facebook and other platforms before ever visiting a publisher's own website.

This creates an obvious tension. Publishers want audiences to visit their properties, while social platforms are designed to keep users inside their own ecosystems for as long as possible. Instagram benefits when users continue scrolling Instagram. TikTok benefits when the next video is watched on TikTok. YouTube operates according to the same basic incentive.

Publishers therefore cannot rely on posting a headline and a link and expecting social audiences to behave like search users. Content has to provide value within the environment where it appears.

A 30- or 60-second video does not need to contain every fact in an investigation. In most cases, it should not. Its purpose may be to capture attention, establish why the subject matters and give the audience enough reason to explore it further.

The full article can provide the evidence, statistics and detail. A podcast can offer a longer hands-free experience. A short vertical video can communicate the central idea quickly, while a longer YouTube or Facebook video can provide explanation, personality and additional visual context.

The formats are therefore not competing versions of the same product. They perform different roles at different stages of the audience relationship.

Converting an article is easy. Adapting it well is harder.

AI has dramatically reduced the technical difficulty of generating audio and video. Creating a synthetic voiceover or producing an automated clip is no longer the hard part. The more important challenge is making the result appropriate for the medium and then getting it in front of the right audience.

A text written for reading is not automatically a good audio script. Written articles often rely on headings, links, charts, formatting, quotations and visual hierarchy. Readers can stop, reread a sentence or scan several paragraphs ahead. Audio is linear, so information has to be structured differently.

The same principle applies to video. An article cannot simply be shortened and placed over stock footage. A TikTok video, an Instagram Reel, a YouTube video and a video embedded on a news website may all originate from the same reporting, but the scripting, pacing, visual language and length should be adapted to the platform.

Even short-form platforms should not be treated as interchangeable. The hook, structure and expectations of a TikTok audience can differ from those of Instagram or YouTube Shorts. The reporting may remain unchanged, but the way the story is introduced and presented should respond to the environment in which it is being consumed.

The goal should therefore be content transformation rather than simple format conversion.

We discuss this in Automatic Audio Generation: Turning Written Articles into Immersive Listening Experiences, where the focus is not simply on turning characters into speech, but on creating an experience that makes sense when consumed as audio.

The same idea applies to video. One piece of journalism can become several different media products, but each product still needs to respect the way its audience consumes that particular format.

Audio quality comes before almost everything else

Publishers beginning to experiment with automated video often focus immediately on visuals, avatars and animation. In practice, poor audio can damage the experience faster than simple visuals.

People will often tolerate a relatively basic visual presentation if the information is useful and the sound is comfortable to listen to. Bad audio is much harder to ignore. Unnatural pacing, inconsistent volume, harsh frequencies or incorrect pronunciation can make even good journalism exhausting to consume.

That is why audio quality should be treated as part of the publishing product rather than as a technical checkbox, including when the final product is video.

In Studio-Quality Narration, Automatically, we look at details such as loudness consistency, sibilance, low-frequency control and other elements that affect whether generated narration feels polished enough for repeated listening.

Voice consistency is another important part of the equation. A publisher already has a visual identity, editorial tone and recognisable style. Audio can become another part of that identity.

Voice Cloning & Brand Consistency explores how a consistent voice can be used across narrated articles, podcasts and other formats without requiring the same person to manually record every piece of content.

For international publishers, language support expands the opportunity even further. MediaThrive currently supports dozens of languages for AI narration and podcast generation, allowing the same editorial infrastructure to serve different audiences without creating an entirely separate production workflow for each market.

More details are available in MediaThrive Now Supports 74 Languages for AI Audio Narration and Podcast Generation.

News and evergreen content perform different jobs

A strong publishing library needs a balance between timely news and evergreen material.

News attracts attention around what is happening now. It responds to current demand, social conversation and breaking developments. Evergreen content works differently. Guides, explainers, analysis and reference articles can continue generating search traffic and audience value long after publication.

Both become more valuable when they are available in multiple formats.

An evergreen article can remain useful for years, which means its audio narration, podcast version or video adaptation can also continue generating consumption. That creates a fundamentally different economic model from a piece of content that is published once and effectively forgotten.

Podcast archives are a good example. With dynamic ad insertion, older episodes can continue carrying current advertising rather than being permanently tied to the campaign that existed when the episode was originally released.

We explain that model in Dynamic Ad Insertion for Podcasts: The Publisher's Guide to Turning Your Audio Archive Into Ad Revenue.

For publishers, the broader lesson is that an archive should not be viewed simply as old content. A strong evergreen library can become an active media asset that continues to generate attention and revenue.

News-focused AI needs to understand what is happening now

There is another important distinction between evergreen publishing and news.

Generative AI is already very capable of producing generic evergreen material. Given the right prompt, a system can create another guide, explainer or list-based article within seconds.

News is fundamentally different. Before producing anything useful, a newsroom needs to know what has happened, whether it matters, how quickly the story is developing, whether reliable sources are confirming it and whether it is relevant to its particular audience.

This makes monitoring an important part of any serious AI-assisted news workflow. Generation should not happen in isolation from what is taking place in the real world.

A more useful process begins by monitoring relevant information sources and identifying meaningful developments. Editorial teams can then verify the information, decide whether the subject deserves coverage and determine the appropriate editorial angle. Only after those decisions have been made does automation become useful for creating or adapting the resulting content.

In other words, the most valuable AI systems for publishers will not simply be systems capable of generating another generic article. They will help newsrooms understand what deserves attention and then reduce the production work required to bring that journalism to audiences across multiple formats.

AI changes the economics of publishing

There is no good reason to claim that AI-generated content will automatically be better than human-created content. In many cases, it will not.

Journalists and editors provide reporting, judgement, context, taste, accountability and an understanding of why a story matters. Those are still fundamental to credible publishing.

What AI changes is the cost of repetitive production work.

A journalist does not need to manually record a narration every time an article is published. An editorial team does not need to create every podcast version from scratch. The same reporting does not need to be manually reformatted every time it moves to another distribution channel.

The most practical model is therefore not human versus AI. It is a workflow in which human editorial judgement determines what deserves to be created, while automation makes it possible to produce and distribute more versions of that work.

When MediaThrive passed 10,000 automatically generated audio articles, the most important aspect of that milestone was not the number itself. It demonstrated what happens when the marginal cost of turning an article into another format becomes very small.

Historically, ten articles meant ten primary media assets. Today, the same ten pieces of journalism can potentially become narrated articles, podcast episodes, short-form videos, longer videos and multiple social assets.

That does not mean publishers should flood every channel with low-quality material. Quality and editorial control still matter. What changes is the number of formats in which a strong piece of journalism can economically exist.

Content recycling should become part of the publishing workflow

One of the biggest opportunities for publishers is therefore not simply creating more articles. It is getting more output from every strong article that already exists.

The reporting has already been completed. Sources have been contacted. Facts have been checked. The article has been edited, formatted and optimised for search. A significant part of the cost of producing the journalism has already been paid.

At that point, limiting the story to a single webpage is an artificial constraint.

A well-designed content recycling workflow can turn the article into an optimised narration for the website, a podcast episode, a short vertical video for TikTok, Instagram or YouTube Shorts, a longer video for YouTube or Facebook and a set of social-native extracts built around statistics, quotes or individual arguments from the original story.

The underlying journalism remains the same, but the packaging changes according to the medium. This distinction matters because successful recycling is not about publishing identical content everywhere. It is about adapting the same editorial asset to the behaviour and expectations of different audiences.

Custom Podcasts by MediaThrive is one example of this principle in practice. Existing editorial content can be turned into podcast episodes and feeds without requiring a separate studio production process for every article.

We also explored the changing relationship between social platforms and podcast discovery in How New Listeners Are Reshaping Podcasting and What It Means for Publishers.

This connection between formats is important. The person who discovers a publisher through a short video today may become a podcast listener tomorrow and a regular reader or subscriber later.

More formats create more monetization opportunities

Every additional format also creates new commercial possibilities.

Audio can support sponsorships, audio advertising, podcast distribution, dynamic ad insertion and premium listening products. Video can create new advertising inventory, sponsorship opportunities and platform-based monetization.

Evergreen content becomes particularly interesting because the value of the asset can continue accumulating long after the original publication date. A useful article that continues attracting audiences can also continue generating listening, viewing and advertising opportunities through the additional formats created from it.

The wider growth of spoken content supports this opportunity. Edison Research has documented strong adoption of podcasts, online audio and audiobooks, showing that audiences are increasingly comfortable consuming substantial amounts of information without looking at a screen.

We explored the audiobook side of that behaviour in Audiobook Reach Hits a New High and It Signals a Bigger Shift in Audio Consumption.

The takeaway for publishers is not that every newsroom should start producing audiobooks. It is that listening has become a mainstream information behaviour. Publishers already own large libraries of material that can participate in that behaviour.

The future is multi-format publishing

The next stage of digital publishing should not be framed as text versus audio or text versus video.

Text remains exceptionally valuable. It is searchable, precise, easy to reference and capable of communicating enormous amounts of detail. SEO will continue to matter because discovery will continue to matter.

But text cannot occupy every moment in a person's day.

Audio creates consumption opportunities when screens are unavailable or inconvenient. Video helps publishers compete for attention on platforms where audiences already spend their time and can act as both a complete media experience and an entry point into deeper journalism. Monitoring helps editorial teams identify which developments deserve attention now, while AI makes the transformation and distribution process economically feasible at a much larger scale.

The publisher's role is therefore becoming broader. It is no longer enough to create a good article and wait for people to find it. Publishers need to understand where their audiences are, how those audiences prefer to consume information in different situations and which format gives each story the best chance of being noticed.

The journalism can still begin with the article. It simply does not have to end there.

A publisher can create the reporting once and then adapt it intelligently for listening, watching, social discovery and deeper reading. The result is not necessarily more journalism. It is more value, more distribution and more opportunities for audiences to engage with the journalism that already exists.

SEO can help win the search result. The next challenge is winning what happens after the click.

Ready to transform your content?

See how MediaThrive's AI-powered platform can help your media team create audio and video content at scale.

No credit card required · 14-day free trial

Stay ahead in AI media

Get weekly insights on media automation, AI content creation, and industry trends.