All articles

Pronunciation should come first because it changes what you hear

1 September 2026

Choc Education


Most English courses treat pronunciation as finishing work. Students learn grammar, memorize thousands of words, and postpone sound training until they feel “advanced enough.”

The order should be reversed.

For Mandarin-speaking learners in Taiwan, pronunciation is often the most productive place to begin because the sound system affects speaking, listening, vocabulary storage, and fluency at the same time. When the brain stores light and night under nearly identical sound patterns, the problem appears during conversation long before the learner opens their mouth. Both words arrive through a blurred channel.

Fix the map first.

A learner who builds accurate sound categories early gains a cleaner mental representation of every new word. That makes the word easier to recognize, retrieve, and pronounce. Thirty minutes spent correcting one recurring contrast can improve hundreds of later encounters with English.

Pronunciation changes what you hear

Listening feels like a receptive skill, yet the listener is constantly making predictions. The brain compares incoming speech with sound patterns already stored in memory. If those stored patterns are inaccurate, even familiar vocabulary can disappear inside a normal sentence.

Consider this exchange:

“Turn right after the light.”

A learner who has an unstable English /r/–/l/ distinction may know all six words on paper. At conversational speed, however, right and light compete for the same perceptual space. The learner hears sound, searches the wrong categories, and falls half a sentence behind.

This explains a common classroom mystery: a student understands the transcript immediately after failing to understand the recording. Vocabulary was never the main obstacle. The written form supplied distinctions that the learner’s ear had not yet learned to detect.

Flege’s Speech Learning Model explains how second-language sounds are filtered through categories formed for the first language. When an English sound seems “close enough” to a familiar Mandarin sound, the brain may assign both to one category. Training must make the difference perceptually meaningful before accurate production becomes reliable.

Pronunciation work therefore belongs inside listening practice. It gives the ear sharper categories.

Why errors become fossilized

Larry Selinker introduced fossilization in his 1972 paper “Interlanguage”. The term describes second-language patterns that become stable even after years of exposure. Learners keep using them because the patterns communicate well enough, receive little correction, and gradually become automatic.

Pronunciation is especially vulnerable. Speech movements happen in fractions of a second. Once the tongue, jaw, and breath system have repeated a pattern ten thousand times, conscious knowledge has limited control over it.

A learner may know that rice and lice differ. During a presentation, attention shifts toward ideas, grammar, slides, and the audience. The older motor habit returns.

Early correction reduces that accumulated resistance. It also prevents a faulty sound representation from attaching itself to every new word. Correcting /r/ after learning 300 words containing /r/ is manageable. Correcting it after 8,000 words requires rebuilding thousands of stored forms.

This connects pronunciation directly to the mental lexicon. Words are stored with sound, meaning, grammar, and associations. A vague sound form creates slower retrieval and weaker listening recognition.

Common trouble spots for Mandarin speakers

Every learner has an individual accent, but several patterns appear frequently among Mandarin-speaking students in Taiwan.

1. The /l/–/n/ merger

Mandarin distinguishes initial l and n, yet some regional varieties and home-language backgrounds weaken or merge the contrast. The two sounds also use nearby tongue positions, which makes confusion easy under time pressure.

Useful pairs include:

  • light / night
  • low / no
  • fly / fine
  • collect / connect

The learner needs to feel the physical difference. For /l/, the tongue tip touches the ridge behind the upper teeth while air escapes around the sides. For /n/, air travels through the nose. Pinching the nose during /n/ creates an immediate, oddly effective diagnostic: the sound stops.

2. English /r/ versus /l/

Mandarin ㄖ and English /r/ are produced differently. English /r/ usually has no tongue-tip contact with the roof of the mouth. English /l/ requires contact near the alveolar ridge.

This distinction changes meaning in pairs such as:

  • rice / lice
  • road / load
  • correct / collect
  • arrive / alive

A mirror helps less than a clear physical instruction. Hold /r/ for two seconds without touching the tongue tip to the roof of the mouth. Then switch to /l/ and make firm contact. The contrast becomes a movement the learner can monitor.

3. Stress-timed rhythm

Mandarin tends toward more even syllable timing. English gives greater prominence to stressed syllables and compresses many unstressed ones. The labels describe tendencies rather than rigid laws, but the practical difference is easy to hear.

A learner may pronounce every word clearly in:

“I’ll send it to you after the meeting.”

Yet equal weight on all eleven syllables makes the sentence harder to process. Natural connected speech emphasizes send, you, after, and meeting, while words such as it, to, and the become shorter.

This affects listening because reduced forms rarely resemble their dictionary pronunciations. A student expecting a fully articulated to may miss /tə/ repeatedly.

4. Minimal pairs that collapse into one category

English contrasts such as ship/sheep, full/fool, cap/cab, and fan/van can be difficult when Mandarin provides no equivalent contrast in the same position.

Minimal pairs expose the exact boundary the learner needs. Random repetition offers weaker feedback because the learner may repeat both words with the same sound and never notice.

A practical correction sequence

The order matters. Begin with perception, connect perception to movement, and then place the corrected sound inside connected speech.

Step 1: Train minimal pairs

Choose one contrast. Avoid working on five at once.

  1. Listen and identify which word you hear.
  2. Use an ABX task: hear A, hear B, then decide whether X matches A or B.
  3. Produce each word while recording yourself.
  4. Place the words in short sentences.
  5. Continue until identification stays above roughly 85 percent across several sessions.

Ten focused minutes works well. A set of twelve pairs is enough for one practice block. Immediate feedback matters more than volume.

Step 2: Add shadowing

Once the contrast is reasonably stable, use a 15- to 30-second recording from a clear speaker. Read the transcript, mark the target sounds, and shadow half a second behind the audio.

Record five attempts. Compare the target consonants and vowels first; ignore minor accent differences. Then repeat without looking at the transcript.

Shadowing links perception, articulation, speed, and memory. It also trains whole sound sequences, which is why it belongs in a broader system of effective English-learning methods.

Step 3: Train stress and rhythm

Mark the stressed words in the same clip. Tap once for each major stress and fit the unstressed syllables between the taps. A metronome set to 72 beats per minute can make the pattern visible to the body.

Then practice three versions:

  • exaggerated stress,
  • natural stress,
  • conversational speed.

Exaggeration is useful during training. It creates a contrast large enough for the learner to hear and feel before refining it.

What research calls good pronunciation

Research by Tracey Derwing and Murray Munro separates accent, intelligibility, and comprehensibility. A speaker may retain a noticeable accent while remaining easy to understand. The practical target is speech that listeners process accurately and with little effort.

Derwing, Munro, and Wiebe also found that pronunciation instruction can improve comprehensibility, especially when teaching includes suprasegmental features such as stress, rhythm, and intonation. Later work by Kazuya Saito and colleagues likewise shows that comprehensibility depends on a combination of segmental accuracy, word stress, rhythm, fluency, grammar, and vocabulary.

Native-like imitation is an unnecessary goal for most adults. Stable contrasts and clear rhythm produce a far greater return.

Why this matters for GEPT and TOEIC Speaking

GEPT and TOEIC Speaking require the learner to produce language under time pressure. The official TOEIC Speaking scoring guide scores read-aloud responses separately for pronunciation and for intonation and stress. Across speaking tasks, unclear sound contrasts and flat rhythm increase the listener’s processing effort even when the grammar is accurate.

Test-day polishing has limited reach because automatic speech habits dominate under pressure. Early pronunciation training gives the learner a dependable base for reading aloud, answering questions, describing information, and presenting an opinion.

The payoff also extends beyond scores. Clearer internal sound categories help learners follow meetings, catch names over the phone, and recognize familiar words in unfamiliar voices.

FAQ

1. Should beginners study pronunciation before grammar?

They should study both from the beginning. A short daily pronunciation block protects new vocabulary from being stored with unstable sound patterns.

2. Can adults still change a fossilized accent?

Yes. Adults usually need focused perception training, explicit mouth-position guidance, frequent recording, and repeated practice in connected speech. Progress is measured in intelligibility and ease of understanding.

3. How long should pronunciation practice take each day?

Ten to twenty minutes is enough when the target is narrow. One contrast practiced carefully produces more useful feedback than an hour of unfocused repetition.

4. Which should come first: sounds or rhythm?

Start with sound contrasts that change words, then move into shadowing and rhythm. Each stage supports the next, and all three should eventually appear in normal conversation.