Speak Reference

Practice Drills Specifically for Reducing Filler Words

Targeted drills can break the motor habit behind filler words.

Reporter · · 11 min read
Cover illustration for “Practice Drills Specifically for Reducing Filler Words”
Speaking Drills · September 21, 2026 · 11 min read · 2,399 words

Filler words feel like a personality quirk, the verbal equivalent of a nervous laugh. Filler words feel like a personality quirk, the verbal equivalent of a nervous laugh, but they're a trained motor habit, built the same way a golf swing or a bad handshake gets built: repetition, reinforced without anyone noticing. They're a trained motor habit, built the same way a golf swing or a bad handshake gets built: repetition, reinforced without anyone noticing. That means the fix is drills, run the way an athlete runs drills, aimed at the specific mechanism producing the "um," not confidence coaching or "just relax."" It's drills, run the way an athlete runs drills, aimed at the specific mechanism producing the "um."

Speech researchers group fillers into three buckets: sounds (um, uh, ah, er, hm), words (like, well, okay, so, basically, really, very), and phrases (you know, to be honest with you, if that makes sense). People reach for them for a handful of reasons that have nothing to do with nerves: buying a beat while the brain hunts for the next word, signaling a topic shift, softening a blunt statement, or just running a groove worn in by years of doing it. Producing fluent speech takes coordination across more than 100 muscles spanning the respiratory, phonatory, and articulatory systems, so a filler is often as much a motor stumble as a mental one. Even fluent speakers hit five to seven disfluencies per 100 words, roughly one every 14 to 20 words. And the gap that makes this worth fixing at all is that the speaker barely notices doing it, while the listener notices every time. That asymmetry is the whole reason awareness has to come before any drill.

What the research says about how much filler is too much

A study out of the University of Kansas, published in the Journal of Applied Behavior Analysis, seeded practice speeches with fillers at four different rates: 0, 2, 5, and 12 per minute, then had audiences rate the speakers. Five per minute barely moved the needle. Twelve per minute tanked ratings across effectiveness, preparedness, and confidence. So there's a real threshold, and it is somewhere between 5 and 12, not at zero.

The damage wasn't evenly spread across filler types, either. Filler sounds (um, uh) hurt ratings more than filler words (like, so), which matters for anyone deciding where to spend practice time first: kill the "um" before worrying about the "like."

Real-world stakes back this up. In telemarketing calls, success rates dropped once filler use crossed 1.3% of total words spoken, a number that turns "sounds unpolished" into "loses the sale." On the other end of the spectrum sits a genuinely wild data point: U.S. presidential inaugural addresses between 1940 and 1996 contained zero filler sounds. Zero. That's not a realistic bar for a Tuesday status update, but it shows what's achievable when a speech gets rehearsed within an inch of its life.

At the level of a single conversation, occasional filled pauses did not significantly hurt how competent, trustworthy, or warm a speaker seemed. One "um" doesn't brand anyone incompetent. Frequency is the enemy, not presence. Which means the training target is control: getting under that 5-per-minute line.

Diagram: The Filler Threshold: Where Ratings Start to Drop. Visualizes: Visualize the filler-rate research from the University of Kansas study, which tested four rates — 0, 2, 5, and 12 fillers per minute — and found that 5 per minute barely moved…

Why awareness comes before any drill, and how to build it fast

Nobody accurately estimates their own filler rate. Asking a room full of professionals how often they say "um" produces answers that are consistently low, often by a lot, because the brain filters its own disfluencies out in real time the same way it ignores the sound of its own breathing. Listeners get no such filter. Closing that gap is the actual first step, before any drill, any pause reflex, any of it.

The baseline exercise is unglamorous: record two to three minutes of talking off the cuff on any topic, then sit through the playback and tally every filler by hand. A meeting platform that auto-generates transcripts turns this into a search-and-count exercise rather than a laborious re-listen; run the transcript, search "um," search "like," search "you know," and count hits.

There's a quick diagnostic for the harder-to-spot filler words: pull the questionable word out of the sentence and read it again. If the sentence loses nothing, that word was dead weight. "So, basically, the numbers were good" reads exactly as clearly as "The numbers were good." That word being safely removable is the tell.

Take that first three-minute count and convert it to fillers per minute. That number is the Week 0 baseline, and it gets checked again at Week 2 and Week 4 against the 5-per-minute target established by the Kansas study. Most people find the first count genuinely shocking, higher than they'd have guessed by a wide margin, and that shock does real work. Once it becomes a number on a page, the problem can go down instead of remaining abstract.

Awareness by itself changes nothing. It's the prerequisite, not the cure. The rest of this piece is where the actual habit gets dismantled.

Drills that target conscious awareness, catching fillers in the act

The Start/Stop Drill puts a second person in the room as a human tripwire. The partner listens and interrupts, out loud, the instant a filler lands, and the speaker restarts from the top of the sentence, or the top of the whole speech, depending on how brutal the version being run is. It is, by design, repetitive and irritating. That's not a flaw to apologize for; that's the mechanism. Repeated interruption trains the speaker to catch the filler impulse earlier each time, and eventually the catch happens before the sound does.

A lighter-weight version swaps the verbal interrupt for a raised hand or a tapped glass every time a filler slips out. One variant on this raises the stakes further: once a signal sounds, the goal becomes avoiding a second one before the exercise ends. It's faster to set up than a full restart drill and works fine as a starting point before graduating to the harsher version.

Mirror recitation strips out the partner requirement. Speaking to a mirror makes visible the physical tells that ride along with a verbal filler, the dropped gaze, the shoulder shift, the hand that suddenly needs something to do, often a half-second before the sound itself. Catching the physical tell early gives a speaker more runway to swap in silence instead.

What all of these share is simple: they take something invisible to the speaker and force it into the open, in real time, with an external signal loud enough to notice. That's the entire job of this category of drill.

Drills that build the pause reflex, replacing the filler with silence

Silence feels endless to the person standing there producing it and reads as composed, even authoritative, to everyone listening. That mismatch is the whole opportunity. A deliberate pause, dropped periodically between phrases, reframes silence as rhythm rather than failure, which is the opposite of how most people experience it mid-sentence.

The 30-Seconds-Without-Filler drill is talking on any topic for 30 to 60 seconds, and every time the urge to say "um" occurs, swapping in dead air instead. No sound allowed, full stop. Start with a topic that requires zero thought, a favorite meal, a weekend plan, so the drill trains the pause reflex without also demanding real cognitive work. Raise the topic difficulty once the reflex starts holding on its own.

The One-Minute Drill runs the same idea with a harder rule: a full 60 seconds, zero fillers, timer running. Begin with subjects known cold, then work up to unfamiliar or complicated ones as the reflex solidifies. Run daily, this drill is what turns "remembering to pause" into something that just happens without a decision being made.

Slowing overall pace helps here too, since it buys more thinking time between words and shrinks the gap that fillers exist to plug. Fillers cluster hardest at the transitions between points. Swap the filler urge for a breath first. The inhale occupies the exact moment a filler would have, at zero cost and with the side benefit of actually helping the voice.

Drills that build fluency under pressure, reducing the cognitive need for fillers

One way to think about fillers is as the brain stalling for time while it searches for the next word, a retrieval delay more than a confidence failure. Whatever the exact mechanism, it points somewhere useful: drills that speed up word retrieval should shrink the window a filler has to fill.

Verbal fluency speed drills do exactly that. Pick a letter, name as many words starting with it as possible in 60 seconds. Or pick a category, animals, cities, kitchen tools, and rattle off as many as come to mind before the clock runs out. No partner needed, no equipment, works as a daily warm-up in a car or a waiting room. The mechanism is speed: training the brain to grab words faster under a small amount of time pressure, the same kind of pressure that produces a filler in live conversation.

Shadowing comes out of psycholinguistics research and asks for something a little strange: play a fluent speaker's audio and repeat their words at the same time, or with a lag of roughly 200 to 500 milliseconds. It forces the articulatory system, the actual muscles doing the talking, to keep pace with someone else's fluent rhythm, and that repeated matching builds the motor patterns that produce filler-free delivery. Any podcast or recorded speech works as the model, and it's a solo drill from start to finish.

Impromptu-topic practice replicates the exact condition where fillers spike hardest: not knowing what's coming. Draw a random subject and talk about it for a set stretch, with someone tracking fillers as they happen, starting with short responses and stretching the time as the skill builds. A close cousin, drawing a random object or topic and talking for one minute, works the same muscle: thinking before speaking rather than thinking while speaking, since the latter is where nearly every filler lives.

Structure does a version of this same job passively. A speaker working from a known shape, beginning-middle-end, or a simple point-reason-example-point sequence, isn't burning cycles on organizing thoughts mid-sentence. Knowing where the next thirty seconds are headed removes the need to stall while the brain finds the road.

Audio recording drills (using playback as the toughest coach in the room)

Diagram: From Awareness to Automaticity: The 8-Week Timeline. Visualizes: Illustrate the improvement timeline described in the article across three distinct milestones: weeks 1–2 (first dip in fillers, awareness effect kicks in), weeks 3–4 (real…

Recording is unforgiving in a way live conversation never is, because there's no forgetting what just got said. The core version: talk for two to three minutes with no script, listen back, tally every filler by hand. Then run the same topic again, this time pausing on purpose the moment a filler impulse occurs, finishing the thought, and continuing. Repeating that stop-pause-continue loop breaks the automatic pattern and forces something else into its place.

Adding a phone camera stacks a visual layer on top of the audio one. Record a two-minute mock conversation or explanation, play it back, mark every filler, then run it again. The visual track can surface physical habits that audio alone doesn't capture.

Transcript review works as a lower-effort variant. A generated transcript turns a search bar into a filler counter, and it reveals something audio review sometimes doesn't: certain filler words appear in writing before a speaker consciously hears themselves saying them out loud.

Then there's the version with no recorder involved. After a meeting or call that mattered, mentally rewind and ask where the fillers landed, and what triggered each one, a tough question, an interruption, a transition between points. No equipment required, and it builds the same awareness muscle in the middle of real stakes rather than a rehearsed drill.

Whichever version gets used, the output should always be the same number: fillers per minute, tracked at Week 0, Week 2, and Week 4, checked against the 5-per-minute threshold the research identifies as the point below which ratings hold steady.

How long improvement takes, and how to structure daily practice

The timeline isn't instant, but it isn't glacial either. Most people notice a first dip in fillers and hesitation within one to two weeks, purely from the awareness effect kicking in. Real gains in confidence and fluency appear around the three-to-four-week mark. Full automaticity, the pause reflex firing without any conscious effort behind it, tends to run around six to eight weeks.

Frequency beats duration here, and it's not close. Ten to fifteen minutes of focused daily practice outperforms a single hour-long session once a week, because short, repeated sessions give the brain room to consolidate the new pattern between reps rather than cramming it all into one sitting.

Treat this like conditioning, not like a lecture to sit through once. Speaking muscles behave like any other muscle: sharp with regular use, sloppy the moment reps stop. A rough weekly shape might run something like this: daily, a One-Minute Drill or a verbal fluency speed round, solo, no scheduling needed; a few times a week, a partner interrupt session or a short audio-recording loop, five minutes is enough; once a week, an impromptu-topic session for pressure reps; and after any conversation that actually mattered, a quick mental debrief on where the fillers showed up. Track one number through all of it: fillers per minute off a three-minute recorded clip. That single figure says whether the work is landing.

Using AI tools to add real-time feedback to the drill cycle

Some of this drilling can now happen with software doing the counting instead of a partner or a stopwatch. AI-based speech analysis tools can transcribe a practice session and flag filler words automatically, cutting out the manual tally that used to eat up half the exercise. Real-time transcription, run during a practice call or a recorded rehearsal, can surface a live filler count instead of one discovered after the fact on playback.

That's a genuine gain in speed and convenience; the actual behavior change still comes from the same place it always has: repetition, discomfort with silence turning... The tool can count faster than a human partner and never gets bored doing it, but the actual behavior change still comes from the same place it always has: repetition, discomfort with silence turning into comfort with silence, and a number tracked over weeks that either drops or doesn't. Software can hand back the data faster. It still can't do the pausing for anyone.

Sources

  1. Verbal Communication Exercises You Can Do Alone | Oompf
  2. We, um, have, like, a problem: excessive use of fillers in scientific speech | Advances in Physiology Education | American Physiological Society
Filed underSpeaking Drills

More in Speaking Drills