Speak Reference
Filler WordsLong read

Practical Techniques for Stopping Mid-Speech Um Habit

Understand why you actually say um, then measure it to fix what matters.

Reporter · · 9 min read
Cover illustration for “Practical Techniques for Stopping Mid-Speech Um Habit”
Filler Words · September 30, 2026 · 9 min read · 2,124 words

Um is not a character defect. It is a retrieval signal, the sound your brain makes while it is still rummaging through the drawer for the right word, and treating it like a bad habit to be shamed out of existence misunderstands what is actually happening at the neural level. Linguist N. J. Enfield, who teaches at the University of Sydney, describes disfluencies as "traffic signals that regulate the flow of social interaction", which is a more useful frame than "verbal tic" because it explains why the sound exists at all: it tells the listener the channel is still open, don't interrupt, more is coming. Primary causes of fillers include nervousness or speaking too quickly, inadequate preparation time, and infrequently used words that are difficult to remember while presenting, and the overuse of filler words often occurs when speakers feel pressured to respond quickly, inserting unnecessary sounds instead of allowing a brief pause. Speak faster than your thoughts can organize themselves, and the gap between mouth and mind widens, forcing the brain to stall the room with a placeholder sound while it catches up. None of this is rare or shameful. Mark Liberman's research, cited in Babbel's language magazine, found "um" and "uh" appearing at a strikingly steady clip in ordinary natural speech, which makes the filler a near-universal feature of unscripted talk rather than evidence of a weak communicator. If the reflex is cognitive and habitual, built through years of repetition, then the fix has to be cognitive and habitual too. Awareness alone will not touch it. Drills will.

When um becomes a real problem

Calibrate before you panic. A study in the Journal of Applied Behavior Analysis by Laske and DiGennaro Reed found that five or fewer disfluencies per minute did not hurt how effective a speaker was perceived to be, and it was only above that rate that judgment started to erode. That's a real number, not a vibe, and it means most people worrying about their um count are worrying about a rate that isn't actually costing them anything.

Above the threshold, though, the damage is not abstract. The Advances in Physiology Education study from Seals and Coppock found that heavy filler use in scientific and professional presentations drags down speaker credibility and actively hurts how well the audience follows the content. That's the cost: not just sounding unpolished, but leaving your listener with less information than you intended to give them.

Research published in Cognition by Lee and Papafragou in 2026 sharpens the picture further by splitting perception into two layers. That study was indexed around August–September 2026, and its global-level finding is that filled pauses did not affect perceived competence, trustworthiness, or warmth. Zoom out to the overall impression, the big-picture read on someone's competence, trustworthiness, warmth, and filled pauses barely register. Zoom into the actual moment, the specific exchange happening in real time, and filled pauses measurably knock down perceived trustworthiness right then and there. Translate that into a job interview or a client pitch: the overall verdict on you as a person might survive some ums just fine, but the specific instant when you're supposed to be closing the deal is exactly where a heavy filler habit bites.

None of this argues for scrubbing every "um" out of existence. Perfectly frictionless speech can read as rehearsed, even evasive, and Rob Drummond's research on spoken language and identity shows fillers doing real social work: managing uncertainty, cushioning disagreement, signaling that a person is actually thinking rather than reciting. Mild disfluency has texture. Erase it completely and you've replaced a person with a teleprompter. The target, then, isn't zero. It's getting under the threshold where fillers start distracting the listener from the content, and that's a considerably lower bar than most anxious speakers assume.

Why counting your ums is the necessary first move

You can't fix what you haven't measured, and almost nobody has an accurate sense of their own filler rate without a recording to check it against. The Seals and Coppock study lays out a specific protocol: record any speech that matters at least three to five times, review each take, count the fillers, then run it again while consciously swapping pauses in for the noise. That loop, record, count, adjust, repeat, is what turns a fuzzy sense of "I think I say um a lot" into an actual number you can move.

The count does more than track progress. Awareness is always the first step toward improvement. Fillers operate below conscious awareness, so willpower alone is not enough to override something so automatic; the Institute of Public Speaking recommends the practice-and-record protocol as the concrete first move. Reviewing a transcript often reveals clusters, three or more fillers packed into a short stretch, and those clusters usually signal a point the speaker hadn't actually thought through. They're a signal that the speaker hit a point they hadn't actually thought through, a conceptual gap disguised as a verbal one. That distinction tells you where to spend your prep time and where to insert a pause.

AI tools have moved into this measurement role fast. Yoodli sits in the professional and enterprise lane, known for post-session breakdowns and, notably, for detecting fillers live during Zoom, Google Meet, or Microsoft Teams calls and nudging the speaker mid-conversation. Those nudges are private, real-time coaching prompts visible only to the speaker, offering subtle cues like slowing down or reducing filler words while the call is still happening. Speeko leans into filler counting, meeting timers, and impromptu-speaking drills, and carries a large library of premium exercises. Toastmasters clubs and members in fact use Speeko's tools for speech practice, meeting timers, filler word counting, and Table Topics-style impromptu speaking. There's also a simpler format gaining traction: a daily single-prompt app that records a short verbal response and hands back a 0-to-100 score, which folds the measurement-and-reps loop into something a person actually does every day rather than saving it for the night before a big talk.

Worth a caveat before anyone treats a dashboard as a cure. Detection accuracy in these tools runs high in controlled lab tests but drops noticeably once you're in a messier, real-world setting, and the more damning finding is that accurate detection showed no correlation with an actual drop in filler frequency after four weeks of use. Spotting an um is not the same skill as un-learning the impulse that produces it. Measurement gets you to the starting line. It doesn't run the race. An analysis found that AI speech coaching tools, having entered this measurement role around 2025–2026, have segmented into deep analysis (post-session), real-time coaching (during a call), and habit-formation (drills). Wellspoken's AI analysis identifies filler rate, pace, clarity, hedging, structure, conciseness, confidence, and pronunciation, and delivers a Wellspoken Index score on a point scale. That Wellspoken Index is specifically a 1000-point scoring system breaking speech down across six dimensions: structure, conciseness, confidence, pronunciation, filler rate, and pace.

The intentional pause: the single most effective replacement for um

Silence feels like falling off a cliff to the person talking and like a completely normal beat to everyone else in the room. That mismatch is the whole game. You cannot say "um" while you are not speaking, so a deliberate pause gives the brain exactly the retrieval time it was trying to buy with the filler, minus the noise. And the pause does something the filler can't: where um broadcasts uncertainty, a clean beat of silence reads as someone choosing their words on purpose, which is close to the opposite signal.

The reason it feels unnatural at first has nothing to do with the listener and everything to do with the speaker's internal clock. Anyone conditioned to believe that silence means losing the floor will instinctively rush to fill any gap, and a pause that sounds completely normal on playback can feel, in the moment, like an eternity. That gap between felt time and actual time has to get walked through in practice, not reasoned away in theory.

The drill is mechanical, and that mechanical quality is why it works. Record a response to a prompt, play it back, and every time an um shows up, re-record just that sentence with a real pause dropped in where the filler used to live. Do that on repeat until the pause stops feeling like a hole and starts feeling like punctuation. Slowing the overall talking pace helps the same cause from another angle: a slower rate gives the brain more runway before it hits a retrieval gap in the first place, so the um reflex simply fires less often, and the calm, deliberate cadence that results is a side effect, not the goal.

Structural preparation as a pre-emptive strike on filler clusters

Pauses fix the moment. They do nothing for the cause. Go back to the cluster idea from earlier: three or more fillers bunched together almost always mark a point where the speaker hadn't decided what came next, a thinking gap wearing a verbal costume. Linguistically, unlexicalized fillers such as "uh" and "um" are non-verbal sounds carrying no specific meaning, while lexicalized fillers like "okay" and "you know" serve more defined functions, and Clark and Fox Tree's 2002 research argued that "uh" signals a short delay while "um" signals a longer one, both read by listeners as searching for a word or holding the floor. No amount of pause practice closes a gap like that, because the problem was never the delivery. It was the thinking.

Structure is the fix, because it hands the brain a path to follow before it ever has to search for one under pressure. Retrieval load drops sharply the moment a speaker knows not just what to say but in what order to say it, and a gap that would otherwise trigger a scramble for words simply never opens. The STAR method, Situation, Task, Action, Result, is the most battle-tested version of this in interview settings, and coaching consensus around it holds that organizing a response into that shape cuts the urge to stall almost by design. The same principle scales past interviews, into pitches, class presentations, tough conversations, anywhere the stakes push a speaker to freeze mid-thought.

Preparation also solves a narrower, more mechanical version of the same problem: some of the most common filler triggers are simply words that don't get used often enough to retrieve smoothly under pressure. Running through key terminology out loud before a presentation shrinks that retrieval lag before it has a chance to show up as a stammer.

The actual drill here is narrower than rehearsing an entire speech front to back. Find the two or three transition points in a planned response where fillers tend to cluster, write out what gets said at each one, and drill those specific bridges rather than the whole script. The joints need the reps. The joints are what need the reps.

Daily reps in low-stakes settings: how the habit rewires

None of this sticks from a single marathon rehearsal the night before something important. The habit that produces um got built over years, through thousands of repetitions of filled speech, and it only yields to a comparable volume of the opposite, deliberate reps of paused, structured speech, done often enough to become the new default. One long practice session doesn't get near that volume.

Short, unscripted speaking under mild social pressure, repeated on a weekly cycle, is what actually rewires the reflex, not polished, fully prepared remarks. Formats built for exactly that exist already: structured impromptu speaking practice is one version, and a daily single-prompt recording with tight scored feedback is another. What both share is the variable that actually matters: the session has to be short enough to fit into an ordinary day and scored precisely enough that the speaker can't fudge whether it's working. One prompt, one recording, one count of fillers, one re-record aiming for a pause instead. Do that daily and the number moves. Do it weekly and it barely twitches.

Treating a dropping filler count as a score to beat, rather than a flaw to manage, changes the entire posture of the practice. Skill-building feels different from self-correction, and it's a lot easier to sustain daily.

Rob Drummond's caution stands as the final word here: the target was never zero fillers. It's fluency that adjusts to context, the ability to slide between the loose, filler-friendly register of a conversation with a friend and the tight, deliberate cadence of a boardroom pitch. A speaker who can code-switch between those two registers on command hasn't erased a flaw. That speaker has actually trained the skill: someone who got lucky on one good day cannot repeat it tomorrow, while someone who has trained the skill can.

Sources

  1. We, um, have, like, a problem: excessive use of fillers in scientific speech | Advances in Physiology Education | American Physiological Society
  2. Um, so, like, do speech disfluencies matter? A parametric evaluation of filler sounds and words | Request PDF
  3. Does disfluency affect social judgments and decisions? Evidence from spontaneous speech - ScienceDirect
  4. It’s, Like, You Know, Science: Why We Use Fillers When We Speak
  5. 10 Ways to Eliminate Filler Words - Institute of Public Speaking
Filed underFiller Words

More in Filler Words