All articles

Pronunciation practice should follow the cost of confusion

12 September 2026

Choc Education


A Taipei project manager has 18 minutes for pronunciation practice before a weekly call, and polishing every unfamiliar English sound would scatter those minutes across dozens of unrelated targets at once. A better syllabus starts with the substitutions that can change many words and make a listener work harder. Pronunciation practice should follow the likely cost of confusion, using actual listener responses to adjust the order.

This article is for an intermediate Taiwanese professional who communicates successfully in English and wants a sharper practice priority than “lose the accent.”

Sound contrasts carry different workloads

The term functional load describes how much work a sound contrast does to keep words apart. Frequency matters. So do the number of word pairs distinguished by the contrast, the positions where it occurs, and variation across English dialects. If replacing one sound with another can turn many common words into competitors, that contrast carries a heavier load. Consider /p/ and /b/. A listener may hear pin as bin, pack as back, or cap as cab. The substitution creates plausible English words, so context must repair the message. The /ð/ and /d/ contrast in then and den separates fewer common pairs, and some English varieties merge it in ordinary speech. Both contrasts can mark an accent. Their cost to comprehension can differ.

This priority fits the distinction made in Choc Education's article on pronunciation changing what learners hear. Listening and speaking share sound categories. A useful category keeps words separate for the speaker and the listener.

What listeners rated

Munro and Derwing tested the functional-load proposal in 2006. Thirteen native English listeners rated 23 sentences produced by Cantonese speakers. The sentences contained naturally produced substitutions classified as high or low in functional load. High-load errors had larger effects on both comprehensibility and accentedness, whereas low-load errors had little effect on comprehensibility in this small set. The original System study called the result exploratory for good reason. Twenty-three sentences and 13 listeners cannot settle a full pronunciation syllabus, and the speakers' Cantonese background limits transfer to Taiwanese learners with different language histories.

In 2022, a study tested error load and accumulation with read-aloud sentences drawn from 20 advanced English users, 11 women and 9 men aged 22 to 36, from several first-language backgrounds. Thirty-one native English listeners rated sentences containing one to four high-load or low-load errors, plus mixed combinations. High-load errors again received poorer comprehensibility ratings. Low-load errors also became costly once more than two accumulated, and ratings for high-load errors deteriorated beyond three. The study in System supports prioritisation while questioning how neatly the middle of Brown's 10-point hierarchy predicts error gravity.

Those samples stay narrow.

Functional load ranks contrasts at the language level. It does not tell a teacher which contrast one learner already controls, how often that learner uses a word, or whether the listener shares the same accent background. Read-aloud sentences also leave out repair, gesture, shared documents and the other resources available on a Microsoft Teams call.

Start with a communication sample

A ranking becomes useful after diagnosis. Record a 67-second explanation that resembles the learner's real task: a product update, a restaurant order, or directions from Banqiao Station. Ask two listeners to write the words they heard and mark any place that required replay. The transcript should preserve substitutions, deleted consonants, misplaced stress and unclear phrasing. Count misunderstandings, rather than every departure from a prestige accent.

Suppose price repeatedly arrives as prize, and cap arrives as cab. Those errors can redirect figures and object names. Suppose the same speaker uses a local vowel quality in coffee, yet every listener identifies the word immediately. The first set earns time sooner. The second can wait unless the learner has a personal reason to change it. Keep the diagnosis narrow. Two listeners may adapt quickly to a familiar colleague, and a pronunciation teacher may understand speech that a new client cannot. A third listener with less exposure to the speaker's accent can change the priority list. My preference is to keep one target for two sessions and collect new listener evidence, rather than rotate through six attractive minimal-pair pages.

Build a 22-minute practice loop

Choose one high-cost contrast found in the sample. Use words the learner actually needs, including names and numbers from the relevant domain. Eight isolated minimal pairs are enough to establish the distinction. Move into phrases quickly because neighbouring sounds and sentence stress alter the signal.

  1. Listen to 8 mixed items and choose the word heard. Show the answer after each response.
  2. Produce 6 words into a phone recorder, with a carrier phrase such as “Please confirm ___.”
  3. Place 4 of those words in short messages that contain a real choice: “Please send the pack” and “Please send it back.”
  4. Give the recordings to a partner who has not seen the script. The partner writes the word and the requested action.

The final round supplies the useful score.

Mouth position, waveform displays and phonetic symbols can help during practice, yet a listener's recovered message decides whether the contrast is working, and confidence records how secure that decision felt. A partner who writes the right word with 51 percent confidence has supplied different evidence from one who answers instantly. Repeat the check 53 hours later with four new words and a different listener. If the contrast holds in isolated words and collapses inside sentences, slow the phrase and inspect stress or consonant clusters. If perception fails before production begins, return to identification. If listeners understand every item, move to the next communication sample.

Leave low-cost features in their place

Learners may choose accent work for identity, professional presentation or pleasure. Those are legitimate goals. Functional load answers a narrower scheduling question: where can limited pronunciation time reduce likely misunderstanding? It cannot rank a person's accent as better or worse. Some communication failures also sit outside segmental pronunciation. Wrong word stress can hide a familiar item. Missing vocabulary can produce a pause that listeners interpret as uncertainty. A vague explanation remains vague with perfect consonants. When a sound contrast produces no transcription errors across three listeners, stop polishing it for the week.

Before the next call, record one 67-second update, circle the first two words that a listener mishears, and spend the 18 minutes on the contrast responsible for the more expensive confusion.