Product design · Concept

Making long podcasts searchable

You heard a two minute story from a three hour episode. Finding it again should not mean listening to the whole thing.

jump to the solution

The problem

Inefficient podcast discovery

Podcast listeners struggle to find specific content inside lengthy episodes, or to judge quickly whether an episode is worth committing to at all. The result is frustration and abandoned listening.

As a user, I want to search for specific conversationsSo that I can find the part I heard elsewhere, like on a clip, without re-listening to the entire episode.
As a user, I want to judge an episode from short clipsSo that I can decide whether the full episode is worth my time before I start it.
Role
Product Design, Visual, Strategy
Timeline
Aug to Dec 2024
Method
Iterative prototyping and usability testing

Social proof

The behaviour was already there, the tooling was not

Listeners who try previews or summaries are twice as likely to finish full episodes (Buzzsprout)
80%Of podcast listeners prefer apps that help them discover shows quickly (Podnews)
35%Growth in short-form audio and video consumption over two years (Deloitte Digital Media Trends)

Job to be done: enable podcast listeners to explore content in a time-efficient and personalised way.

Approach

Three capabilities, and a long episode stops being one block

The whole idea rests on treating an episode as addressable rather than linear. Three features do that, and each one answers a different reason a listener gives up: they cannot find a part they half-remember, they cannot tell if the episode is worth an hour, and they cannot skip to the subject they came for.

The episode transcript with a passage highlighted and a search control beneath it.
In-episode search. The transcript is the index. Type a phrase you half-remember and the player jumps to the second it was said, instead of you scrubbing for it.
The Clips tab, a vertical feed of short excerpts from full episodes.
Clips and highlights. Short excerpts pulled from full episodes and stacked into a feed. Sampling replaces committing, which is the only honest way to decide whether an hour is worth it.
The chapters tab listing named sections of an episode with timestamps.
Topic segmentation. Named chapters with timestamps, so a three-hour episode becomes six things you can enter at rather than one you have to start.

Primary user flow

Every capability one step from home

Before any visual work I mapped the flow end to end. The test was simple: if a feature could not be reached from the home screen in one move, it would not get used, because the listener who needs it does not yet know it exists.

Flow diagram from app launch through onboarding to the home screen and its five branches.
The flow. Launch, walkthrough, sign-in, then one question that does the onboarding work: pick your genres and creators. Home branches five ways, and Search, Clips and Explore are three of them.

Wireframing

Settling structure before surface

Early wireframes tested layout and hierarchy across onboarding, home, the episode view and the clip. They went through several rounds, and the question at every round was whether the structure still held without colour or type doing the work.

Hand-drawn wireframes for onboarding, home, the episode view and the clip card.
Wireframes. Onboarding across the top, then home and the clip screens worked out in red. The clip card was already the hardest element here, and it stayed the hardest all the way through.

Visual design ideation

Five clip cards, tested rather than argued about

The clip card carries the entire sampling idea, and it has four jobs competing for one small surface: the pulled quote, which episode it came from, how long the full thing is, and the way in. I built five directions and put them in front of users rather than picking the one I liked.

Clip card with the play action inside the quote block and the source in a dark strip below.
One. Full Episode sits inside the quote block; the source runs along the bottom as a dark strip. The action reads clearly, but it interrupts the quote it is meant to serve.
Clip card with cover art moved up beside the title and the play action in the footer.
Two. Cover art promoted into the header beside the title, action moved down to the footer. Strong sense of source, but the header now carries art, title, duration and an overflow menu.
Clip card with an Episode Details label and the play action on its own row.
Three. An explicit "Episode Details" label above the action. Unambiguous, and a row of chrome between the quote and its attribution.
Clip card with a progress bar under the header and the duration shown as a pill in the footer.
Four. Progress bar under the header, duration reduced to a pill. Quietest of the four, but the source still sits in the same dark field as the quote and blurs into it.
The selected clip card, with a progress bar and the episode shown on a light card.
Selected. The one users chose. The quote owns the card at full contrast; the source lifts onto a light card so it reads as a different kind of thing; the duration doubles as the play action. A progress bar under the header says how far through the clip you are without adding a control.

The winner is not the prettiest of the five. It is the one where a listener could tell, without reading, which part was the quote and which part was the podcast.

Final prototype

The system, end to end

Onboarding, browse, clips, search, player and profile, built out as a working high fidelity prototype. One dark ground, one accent, and type carrying the hierarchy, so the artwork inside a clip is the only colour competing for attention.

Grid of every Vocali screen, from onboarding through to the player.
Every screen. Reading across: onboarding and genre selection, browse and library, the clips feed, search results, and the player in its chapter, transcript and related states.

The mark came last and deliberately stayed crude: an ear and a drip, drawn in one weight so it survives at sixteen pixels.

Audio products tend to reach for waveforms. A waveform says sound. It does not say that this one is about pulling something out of a longer thing, which is the only claim the product actually makes.

The Vocali mark, an ear shape with a drip beneath it.
The mark. One colour, one weight, no gradient. It had to sit on a dark player without a container.

Reflection

What this project taught me

Vocali was the first time I designed an interface whose core value came from a machine reading the content, generating clips, segmenting topics, making audio searchable. The design question was not what the model could do. It was how much of its output a listener should be shown, and how confident the interface should sound about it.

That question, how to present generated results honestly, with clear handling for loading, empty and wrong, turned out to be the same one I now work on every day in agent interfaces. Vocali is where I first hit it.

Next project
Aster UI, an agentic interface component library