A 29-year-old semiconductor engineer in Hsinchu can score 12 out of 12 on a ship–sheep exercise recorded by one teacher, then miss the same vowel contrast when a customer speaks on Microsoft Teams. The first score may reflect a memory for that teacher's voice. English ear training should change the speaker, word and sentence before it declares a sound learned. Variation turns a tidy listening score into a test of transfer.
This article is for an intermediate learner who reads workplace English comfortably and still loses familiar words when the voice changes.
A sound arrives inside a voice
Vowels and consonants never reach the ear alone. Their acoustic shape shifts with the speaker's vocal tract, pitch, speed, accent and neighbouring sounds. The /l/ in light occupies a different position from the /l/ in collect. A learner may discover a cue that separates two buttons in one exercise while relying on details tied to a single recording. That narrow learning can look convincing. Accuracy rises because the same talker and word list return. A new colleague speaks, and the cue disappears. The teaching problem is generalisation: can a distinction learned from one set of examples guide perception of an untrained word or voice? A pronunciation activity should include an answer key during practice and reserve unfamiliar recordings for the test. Repeating the training file measures memory for the file.
Choc's article on pronunciation changing what learners hear explains how speech categories affect listening. Voice changes test those categories.
Six listeners and 68 word pairs
Logan, Lively and Pisoni's 1991 experiment trained six Japanese listeners to identify English /r/ and /l/. They used 68 minimal pairs spoken by five talkers across 15 training sessions, with the target sounds appearing in several positions. Participants received immediate feedback. Their identification improved from pretest to post-test, and the researchers also tested unfamiliar words and a new talker. The original Journal of the Acoustical Society of America paper reported small, reliable gains. The design bundled several ingredients: natural recordings, multiple voices, varied phonetic positions, an identification task and feedback. Six people are far too few to assign the gain to one ingredient with confidence.
The 1993 follow-up compared multi-talker and single-talker training, finding that listeners in the multiple-voice experiment improved and transferred learning to new words from both a familiar speaker and an unfamiliar speaker. The single-voice group improved on trained material yet failed to transfer reliably to a new talker. The follow-up paper made talker variability a major part of phonetic training research. These studies concern Japanese learners and /r/–/l/. A Taiwanese learner working on ship–sheep, final consonants or weak vowels presents another language history and another acoustic problem. The classroom principle has to stay modest: change examples and inspect transfer.
A large replication narrowed the claim
Brekelmans and colleagues revisited the same /r/–/l/ question with 166 Japanese listeners in a registered replication published in 2022. Learners in both the multiple-talker and single-talker conditions improved during training. The study found no clear learning advantage for high variability, and evidence for better generalisation to new voices remained ambiguous. If a multiple-voice advantage exists, it was smaller than the team could detect in this sample. The paper, materials and analysis are available through the Journal of Memory and Language. That result changes the lesson plan. Several voices remain useful because learners will meet several voices and teachers need a transfer test. The evidence cannot support a fixed rule that five speakers always teach better than one. Training amount, sound contrast, feedback, voice order and individual aptitude can alter the result.
Build a transfer ladder
One voice may be useful at the start.
Learners who cannot yet hear any difference between two sounds may need 14 slow, clear trials from the same speaker before variation arrives, once they have a plausible cue to follow. I would move from clean to mixed recordings within the same lesson, rather than open with eight accents and call the confusion desirable difficulty. A 19-minute activity can separate practice from transfer. Choose one contrast that causes a real comprehension problem, such as ship–sheep, price–prize or fifteen–fifty. Collect 12 short items from four speakers. Keep recording level similar so volume gives no accidental answer.
Use the items in four rounds:
- Two words from one clear speaker, with the written choices visible and immediate feedback.
- Six words from three speakers, mixed in a new order. Learners identify each item and mark confidence from 1 to 4.
- Four short phrases from an unfamiliar speaker, with no text. A wrong answer triggers one replay and then a transcript check.
- A delayed check 47 hours later uses new words and one new voice.
Confidence exposes lucky guesses.
Someone who is 90 percent sure and wrong needs a different cue from a person who chose randomly. For ship–sheep, the teacher might direct attention to vowel quality and duration. For price–prize, the useful evidence may include the voicing of the final consonant and the length of the preceding vowel. The cue depends on the contrast. Record results in a small grid with columns for familiar voice, new voice, familiar word and new word. If accuracy falls only with a new speaker, add voices. If every speaker causes trouble on final consonants, return to the sound cue. If isolated words are secure and phrases collapse, connected speech or attention load has become the next target.
Keep difficulty attached to a purpose
Changing voices every trial can overload a beginner. Poor microphones and heavy background noise also add variation that teaches little. Start with clear natural speech, then introduce the range the learner is likely to meet: two colleagues from Singapore, a manager from California and a British training video, for example. Accent diversity belongs in the course when it matches actual listening demands. Perception work also has limits. A missed word may come from weak vocabulary coverage, an unfamiliar collocation or divided attention during a screen share. Sound training handles one bottleneck. The diagnostic grid should reveal when the bottleneck has moved.
Tomorrow, keep two items from your last file, replace ten with new words from four speakers, and save a fifth voice for the 47-hour check.