
TL;DR
Eliminating every filler word is the wrong goal — controlled fluency is the right one. Research found that filler rates around 12 per minute significantly hurt speaker perceptions, while roughly 5 disfluencies per minute didn't damage perceived effectiveness, giving you a practical training reference instead of an impossible standard. Fillers appear when your mouth outruns your planning, driven by planning gaps, cognitive load, nervousness, turn-holding, and second-language processing. Fix it with a narrow feedback loop: record a two-minute answer, mark each filler, then repeat the same prompt targeting only silence — one study found immediate feedback cut filler use by more than half without increasing anxiety. Replace fillers with silent pauses or structural bridges ("I'd break that into two parts"), and pace answers using semantic chunking with pauses between context, problem, action, result, and lesson. Use memory cues rather than scripts, since a memorized answer produces more hesitation when the question changes wording. Track one behavior at a time, and judge success by whether fillers distracted from your message — not by a perfect transcript.
Most advice on how to stop saying filler words starts with the wrong instruction: eliminate every “um,” “uh,” “like,” and “you know.” That standard sounds disciplined, but it can make interview practice more stressful and your answers less natural. The practical target is controlled fluency, meaning you can pause, plan, and continue without allowing verbal placeholders to crowd out your ideas.
A small amount of disfluency is normal in spontaneous speech. The interview skill isn't sounding machine-perfect. It's keeping your delivery clear enough that the listener focuses on your judgment, experience, and results.
How Do You Stop Saying Filler Words?
Most advice starts with the wrong instruction: eliminate every "um," "uh," "like," and "you know." The practical target is controlled fluency — being able to pause, plan, and continue without verbal placeholders crowding out your ideas.
The research supports a more useful standard than "never hesitate." A 2024 peer-reviewed study found that when filler sounds rose to 12 per minute, perceptions across most categories dropped significantly — but a low nonzero rate of 5 disfluencies per minute didn't hurt perceived effectiveness. For context, typical spontaneous speech contains about 6 disfluent words per 100 words, and natural conversation is commonly estimated at 5% to 10% disfluency overall.
Step 1 — Understand the mechanism. Fillers appear when your mouth moves faster than your planning. The main interview triggers are planning gaps (you know the experience but not which detail goes first), cognitive load (multi-part questions), nervousness (which accelerates pace and leaves less retrieval time), turn-holding (signaling you're not finished), and second-language processing.
Step 2 — Build a narrow feedback loop. Record a two-minute answer, have a partner or tool mark each filler as it occurs, then repeat the same prompt with one goal only — replacing verbal pauses with silence. In one public-speaking study, speakers who received immediate feedback used less than half as many fillers as the no-feedback group, with no adverse impact on anxiety or self-perceived competence.
Step 3 — Replace, don't suppress. When you feel "um" arriving, close your mouth and breathe in silence. If you need a bridge, use a phrase that adds structure: "I'd break that into two parts," "The key point is how I handled the disagreement," or "The example that comes to mind is…"
Step 4 — Pace with semantic chunking. Divide answers into meaningful units — context, problem, action, result, lesson — and pause between them. Take a quiet breath before your first sentence and open with a direct frame like "I'll use one project example."
Step 5 — Keep one correction active per round. Monitoring posture, eye contact, vocabulary, pace, and filler count simultaneously overloads the same planning system that produces the habit.
Why Zero Filler Words Is the Wrong Goal
Great speakers don't necessarily use zero fillers. They manage them.
A 2024 peer-reviewed study indexed on PubMed found that filler sounds measurably changed how speakers were judged. When filler sounds rose to 12 per minute, perceptions across most categories dropped significantly. However, a low nonzero rate of 5 disfluencies per minute didn't hurt perceived effectiveness in the study's listener judgments. You can review the findings in the PubMed study on filler sounds and speaker judgments.
That gives interview coaching a more useful standard than “never hesitate.” Keeping speech at or below about 5 disfluencies per minute appears consistent with maintaining credibility and effectiveness, based on that evidence. It isn't a universal pass or fail line, but it's a practical training reference. You're trying to prevent distracting overuse, not erase every trace of natural thinking.
Why perfection can make you sound worse
Interview answers require retrieval. You're selecting an example, organizing it, choosing accurate language, and monitoring the interviewer's reaction at the same time. If you panic whenever a filler appears, your attention shifts from the answer to self-surveillance.
That can create a familiar cycle:
- You notice an “um.”
- You criticize yourself internally.
- You rush to avoid another pause.
- Your breathing and planning deteriorate.
- More fillers appear.
A completely memorized script can cause a similar problem. If the interviewer changes the wording of a question, you may lose your place and produce more hesitation than you would with a flexible outline. Candidates speaking in a second language can also sound less authentic when they force themselves to imitate an unnaturally fast or polished delivery.
Practical rule: Replace filler-heavy speech with deliberate pauses, not with frantic silence and rigid memorization.
Controlled fluency leaves room for a short pause before a difficult point. It lets you say “I'd break that into two parts” when you need structure. It preserves your natural voice while reducing the sounds that repeatedly interrupt your message.
The listener is evaluating the whole performance, not a transcript stripped of every hesitation. Your examples, reasoning, specificity, and ability to answer the question still matter. Filler reduction should support those qualities, never replace them.
What Actually Causes Filler Words in Speech
Filler words usually appear when your mouth moves faster than your planning. You reach the end of one idea, need time to retrieve the next phrase, and insert “um,” “uh,” “so,” or “you know” to hold the space.
They can also serve useful conversational functions. Research reviews describe fillers as tools for turn-holding, self-correction, pragmatic softening, and buying thinking time. The 2022 review of filler words in spoken communication notes that filler use is connected with nervousness, speaking speed, and preparation gaps. Treating every filler as a character flaw misses the mechanism you need to change.

The main triggers in interviews
Planning gaps are common during behavioral questions. You may know the experience but not yet know which detail belongs first. A filler buys time while you decide whether to start with the situation, your action, or the result.
Cognitive load rises when a question combines several demands. “Tell me about a difficult stakeholder and how you handled the situation” requires you to retrieve an event, identify the conflict, explain your behavior, and describe the outcome. Without a simple answer structure, fillers occupy the gaps.
Nervousness changes breathing and pace. Many candidates accelerate because silence feels exposed. The faster they speak, the less time they leave for retrieval, and the more verbal placeholders they produce.
Turn-holding matters in live conversation. A filler can signal that you haven't finished speaking, especially when an interviewer reacts with a nod, a brief interruption, or an ambiguous pause.
Second-language processing adds another layer. Multilingual candidates may be translating, searching for idiomatic phrasing, or checking grammar while answering. A 2025 study of Chinese English majors connected filler use with L1 transfer, task complexity, and affective factors such as anxiety. Some learners also viewed fillers as making speech sound more natural, so the right aim is reduction where fillers distract, not suppression of authentic communication. The 2025 study of fillers and hesitations in English majors' spontaneous speech provides that context.
Spontaneous speech gives you a realistic baseline. One synthesis reports about 6 disfluent words per 100 words in typical spontaneous speech, while filled pauses alone range from roughly 1.3 to 4.4 per 100 words, depending on the corpus. Natural conversation is also commonly estimated at about 5% to 10% disfluency overall, according to the research synthesis on fillers and disfluency. Those figures don't excuse distracting interview delivery, but they do make shame a poor training strategy.
Awareness Drills and Replacement Techniques
You can't change a speech pattern you can't hear. Start by recording a realistic answer to a question such as, “Tell me about a time you disagreed with a teammate.” Don't restart when you stumble. Let the answer run, then listen once for content and again only for verbal habits.
Write down each filler, including sounds and repeated phrases. Track whether it appears at the beginning of an answer, before a technical term, after an interviewer interruption, or while you're searching for a result. Your pattern matters more than a generic list of forbidden words.
Use a narrow feedback loop
Immediate feedback works better than vague delayed criticism. In a public-speaking study, speakers who received immediate feedback used less than half as many filler words as the no-feedback group during initial speech exposures, and the effect persisted when feedback continued across an entire course. The PubMed research on immediate feedback and filler reduction also reported no adverse impact on state or trait anxiety or on self-perceived communication competence.
Set up the drill like this:
- Record a two-minute answer to one behavioral question.
- Have a coach, practice partner, or speech tool mark each “um,” “uh,” or target phrase as it occurs.
- Repeat the same prompt with one goal only, replacing verbal pauses with silence.
- Compare the first and second attempts, then choose the next target.
The practical repetition target from the study's application is to keep repeating until the count drops by at least half from the first attempt. That number is tied to the specific practice example described in the Dayton study resource on filler-word feedback, not a promise that every speaker will improve at the same pace.
Replace, don't merely suppress
When you feel “um” arriving, close your mouth and breathe in silence. A silent pause usually feels longer to you than it sounds to the interviewer. If you need a bridge, use a phrase that adds structure:
- For a behavioral answer: “The key point is how I handled the disagreement.”
- For a complex question: “I'd break that into two parts.”
- For a result: “The outcome was different from the initial expectation.”
- For retrieval: “The example that comes to mind is…”
These phrases work because they organize the answer rather than disguising uncertainty. The research on “uh” and “um” in spoken language describes these fillers as common floor-holding devices, which is why a planned pause or meaningful bridge is more effective than ordering yourself to “stop.”
For realistic repetition, use Qcard interview practice to rehearse questions without memorizing full scripts. Keep one correction active per round. If you simultaneously monitor posture, eye contact, vocabulary, pace, and filler count, you'll overload the same planning system that produces the habit.

Pacing and Breathing Strategies for Clearer Delivery
Rushing creates the conditions fillers need. Your ideas arrive in chunks, but your speech tries to deliver them as one continuous stream. By the time you reach the middle of a sentence, you're short on breath and ahead of your own planning.
Start with semantic chunking. Divide an answer into meaningful units, then pause between them. For a question about a failed project, your internal map might be:
- context
- problem
- action
- result
- lesson
You don't need to announce every label. You need to let each unit land before retrieving the next one.
Build pauses into the answer
Take a quiet breath before your first sentence, especially after the interviewer finishes speaking. Begin with a direct frame such as, “I'll use one project example.” That opening gives you a clear route and prevents the first few seconds from filling with “so” and “like.”
Consider this rough answer:
“Um, so, there was this project where, you know, we had a deadline, and I kind of realized that the requirements weren't really clear, so I talked to the team and we sort of changed the process.”
The content is useful, but the speaker is trying to carry context, diagnosis, action, and process change in one breath. A paced version is easier to follow:
“I'll use a project with an unclear deadline. The requirements were changing, so I documented the open decisions and reviewed them with the team. That gave us a shared delivery plan, and I learned to confirm decision ownership before work begins.”
The second answer doesn't require a theatrical speaking style. It uses shorter units, explicit verbs, and pauses at points where the meaning naturally changes.
Adapt the method to your processing needs
Multilingual candidates may need more response-initiation time to formulate precise English. Neurodivergent candidates may also benefit from a visible structure, a longer planning pause, or permission to ask for clarification. Those supports aren't evidence of weak communication. They're ways to reduce avoidable cognitive load.
Try a short response frame:
- “The situation was…”
- “My responsibility was…”
- “I took three actions…”
- “The result was…”
- “What I learned was…”
You can write those cues on a card, keep them beside your practice screen, or rehearse them until the structure becomes familiar. Don't script every sentence. Memory cues preserve flexibility, while word-for-word memorization makes recovery harder when the question changes.
A slower response isn't automatically a poor response. If you need time, say, “Let me choose the most relevant example,” then take the pause. That sounds more controlled than filling the gap while your brain searches.
Using AI Tools for Real-Time Filler Word Feedback
Self-recording is valuable because it reveals patterns you can't feel while speaking. Its weakness is timing. If you review a recording later, you may understand that you used too many fillers without learning exactly when the habit began or what triggered it.
Human coaching adds interpretation. A coach can tell you that your fillers cluster before metrics, challenge you with follow-up questions, and distinguish a useful planning phrase from distracting repetition. The trade-off is access, consistency, and cost. A partner may also notice content and confidence while missing every small “uh.”
Real-time AI feedback fills a different role. It can flag filler occurrences as they happen, show whether usage rises during difficult answers, and help you repeat the same prompt immediately. The value comes from a tight loop: speak, notice, adjust, repeat.

What to look for in a feedback tool
Prioritize features that support interview behavior rather than generic speech scoring:
- Live filler detection: The tool should identify recurring forms such as “um,” “uh,” “like,” “basically,” “you know,” “kind of,” “sort of,” and “right.”
- Pacing visibility: Filler counts are easier to interpret when you can see whether rushing or long, unstructured answers are involved.
- Answer-length feedback: A candidate may reduce fillers but still bury the point in an overlong response.
- Follow-up pressure: Practice should include unexpected probing, because controlled fluency has to survive interruption and retrieval demands.
- Memory cues instead of scripts: Resume-grounded prompts can remind you of achievements without forcing you to recite invented wording.
- Privacy controls: Check whether sessions are recorded, how transcripts are handled, and whether your information is used beyond the practice session.
Qcard is one example of an AI interview copilot that surfaces high-level, resume-grounded talking points in real time and includes an Interview Coach for pacing, filler words, and answer length. Its practice and AI mock interview tools are designed for spoken rehearsal, follow-up questions, and delivery feedback rather than word-for-word scripts.
AI shouldn't replace judgment. A live tool can over-alert, misinterpret a purposeful discourse marker, or encourage you to chase a score instead of answering the question. Use it to identify patterns, then listen to the recording yourself and decide whether the filler actually distracted from meaning.
For candidates with anxiety, multilingual processing demands, or neurodivergent communication styles, the best setup is corrective without being punitive. Track one behavior, use short practice rounds, and stop when your attention becomes too divided to produce a meaningful answer.
Building Your Filler Word Reduction Practice Plan
A useful practice plan measures behavior without turning your voice into a spreadsheet. Start by recording answers to common questions and note fillers per minute, where they appear, the answer's approximate length, and whether the response covered the question directly. Use the same question again after a correction round so you're comparing like with like.
A simple rehearsal cycle
Use this cycle across a two-to-four-week preparation window:
- Baseline: Record several realistic answers without correcting yourself.
- Pattern audit: Identify your most common filler and the situations that trigger it.
- Single-target practice: Replace that filler with silence or one bridging phrase.
- Structure practice: Answer from resume-grounded cues, not a complete script.
- Pressure test: Add follow-up questions and practice recovering after an interruption.
- Interview simulation: Rehearse with normal eye contact, camera placement, and answer pacing.
A practical milestone is not “I never say um.” It's that your filler rate approaches the controlled range discussed earlier, your pauses feel intentional, and you can still explain your experience without sounding rehearsed. If your starting rate is high, focus first on reducing repeated fillers during answer openings and before key results. If your interview is close, prioritize stable delivery over ambitious transformation.
Keep the tracking humane
After each answer, ask three questions:
- Did I answer the question asked?
- Did my main action or insight come through clearly?
- Did fillers distract from the message?
That final question matters. A candidate who uses a small number of fillers but gives a precise, relevant answer may perform better than someone who sounds perfectly smooth while avoiding specifics. Review the Qcard interview preparation guide for a broader rehearsal structure, then adapt it to the role and your processing needs.
On interview day, keep only a few cues visible: the question's core demand, the example you want to use, and the result you need to mention. Take a breath before answering. Pause when you need to think. Speak to the interviewer, not to an imaginary filler counter.
The strongest delivery is controlled, flexible, and recognizably yours.
Key Takeaways
- Zero fillers is the wrong target and can make delivery worse — panicking at every "um" shifts attention from your answer to self-surveillance, creating a cycle where you rush to avoid pauses, breathing and planning deteriorate, and more fillers appear, while research shows about 5 disfluencies per minute doesn't harm perceived effectiveness.
- Fillers are a planning problem, not a character flaw — they appear when retrieval lags behind speech, which is why the fix is buying planning time through silent pauses and semantic chunking rather than ordering yourself to stop, and why research describes fillers as legitimate tools for turn-holding, self-correction, and buying thinking time.
- Immediate feedback beats delayed review by a wide margin — a public-speaking study found speakers receiving real-time feedback used less than half as many fillers as the no-feedback group, with no adverse impact on state or trait anxiety, which is why a tight loop (speak, notice, adjust, repeat the same prompt) works better than listening to a recording days later.
- Structural bridges outperform suppression because they organize the answer rather than disguising uncertainty — "I'd break that into two parts," "The key point is how I handled the disagreement," or "Let me choose the most relevant example" all sound more controlled than filling a gap while your brain searches.
- Memorized scripts increase hesitation rather than reducing it — when an interviewer rephrases a question, a word-for-word script leaves you without a recovery path, while a five-part cue frame (situation, responsibility, actions, result, lesson) preserves flexibility, which matters especially for multilingual candidates who need response-initiation time and neurodivergent candidates who benefit from visible structure.
Qcard gives you resume-grounded memory cues, AI-scored practice, mock interviews with follow-ups, and real-time coaching for filler words, pacing, and answer length without forcing you into a script. Visit Qcard to practice turning anxious pauses into clear, authentic interview answers.
Ready to ace your next interview?
Qcard's AI interview copilot helps you prepare with personalized practice and real-time support.
Try Qcard Free