すべての記事

Show learners what the mouth is doing

2026年10月3日

Choc Education


“Open your mouth more” gives a learner a direction with no destination. No target. The instruction leaves the amount and moving part unspecified. A pronunciation model becomes easier to copy when the learner can see the lip opening, jaw position or tongue target, compare it with a physical movement, record the result and receive a precise response from someone who can hear whether the intended word survived.

Pronunciation teaching should use visual articulatory information for sounds that resist listening and imitation. A mirror, a side-view video, a simple mouth diagram or an acoustic display can make part of the movement observable. The visual cue needs sound and teacher feedback beside it.

A sound is also a movement

Learners usually meet pronunciation as symbols and audio. The International Phonetic Alphabet distinguishes categories cleanly, yet a symbol does not tell a learner what their own tongue just did. Audio provides the result after several articulators have moved together.

For a Mandarin-speaking learner working on English /ɛ/ and /æ/, “said” and “sad” may sound too close at first. The production difference includes tongue height, jaw opening and surrounding consonants. A front-facing mouth video shows some jaw and lip information. It cannot show the tongue body clearly.

This limitation should shape the tool. A face or hand gesture can demonstrate visible lip aperture. Hidden tongue movement may need a side diagram, ultrasound, electromagnetic articulography or a teacher's concise verbal cue. One picture cannot reveal the whole vocal tract. The missing view matters during correction.

What the experiments show

Xi, Li, Baills and Prieto studied 99 Catalan-Spanish bilingual university students learning the English /æ/ and /ʌ/ contrast. Their 2024 Language Learning paper compared instruction with gestures representing lip or tongue features against instruction without those gestures. The benefits differed according to what the gesture represented and how the researchers tested pronunciation. The study's practical lesson is specific: visual encoding works better when its relation to the articulatory feature is clear.

An earlier study tested audiovisual training for the English /l/ and /r/ contrast with 62 Japanese learners. Learners received ten sessions using audio alone, natural audiovisual speech or a synthetic talking face. The Speech Communication paper examined perception and production, including transfer to a new contrast. It offers evidence for combining visible speech with sound, while also showing that a face on screen is an instructional condition, not magic.

Real-time biofeedback goes further. Suemitsu and colleagues used electromagnetic articulography to display tongue position while Japanese learners practised American English /æ/. In their 2015 experiment, the groups receiving visual articulatory training improved production, with 21 participants divided across three conditions. That sample is small, the method requires specialised equipment and the target was one vowel. Do not equate a bathroom mirror with laboratory tongue tracking.

Together, the studies support visual cues alongside sound and feedback.

Pick a feature the learner can control

Imagine a 24-year-old graduate student in Taipei preparing an English conference talk. Record one sentence and have a teacher identify the contrast that most affects comprehensibility. Suppose bad repeatedly approaches bed. The lesson target is the /æ/ vowel in stressed words, not the learner's accent as a whole.

Start with perception. Use two speakers and a short identification set: bad/bed, man/men, land/lend and sat/set. Eight tokens can reveal whether the learner hears the category reliably. Keep it short. Our article on changing the voice in ear training explains why one model voice gives too narrow a test.

Then show the movement. For /æ/, use a clear front and side view, mark the greater jaw opening, and give one modest cue for the tongue. Let the learner try the vowel in isolation while watching a mirror. Move quickly into bad, then “a bad plan,” then the original conference sentence.

Avoid a pile of instructions. If the learner is monitoring lips, jaw, tongue, stress and spelling in one attempt, attention fractures. Choose one visible feature for three or four productions. Record the fourth. Compare it with the model and select the next adjustment.

Visual feedback needs a target

A waveform looks scientific and often tells a beginner almost nothing. Even a spectrogram needs a defined question. The teacher may check vowel duration, the first two formants, aspiration or sentence pitch. The display earns its place when the learner knows which trace corresponds to the movement they are changing.

Praat can display vowel formants, and pitch software can show intonation. Those measurements require competent interpretation. Phone speech-to-text provides a different kind of evidence: whether a system recognised the intended word. Recognition errors may come from the microphone, language model or surrounding sentence. Repeat the trial, because one failed transcription cannot diagnose tongue position.

Mirrors have the opposite trade-off. They are cheap and immediate, but they show only the front of the mouth. A learner may exaggerate visible lip movement while the hidden tongue target remains unchanged. Teacher listening and a later intelligibility task keep the mirror from becoming the goal.

Keep the sound inside speech

Articulatory practice often begins with an isolated sound because the movement is easier to feel there. It must soon return to words and sentences. Coarticulation changes every target: the /æ/ in man occurs between two nasal consonants, while bad places it before a voiced stop. A learner who produces a clean vowel alone may lose it at normal speaking speed.

Use a short ladder:

  • identify the sound from two voices,
  • copy the movement in a mirror,
  • say four high-value words,
  • use two phrases under light time pressure,
  • return to the learner's own sentence.

Returning to the learner's sentence protects the purpose of the lesson. If the conference audience can distinguish “men” from “man” in the data description, the adjustment has worked. Matching a native acoustic average is a different goal and usually an unnecessary one. Pronunciation practice should follow the cost of confusion.

Cases where another route is better

Visual articulation is a poor first tool when the main problem is rhythm, misplaced prominence or pausing. A hand movement tracing sentence stress may fit those features better than a mouth diagram. Learners with visual impairment need tactile, auditory or verbal alternatives. Some people become tense when watching their face, and excess tension can worsen production.

Teachers also need to distinguish perception from motor control. If the learner cannot hear which word a speaker produced, production feedback alone leaves the category unstable. If perception is accurate across speakers and production remains unclear, articulatory work deserves more time.

One preference is worth stating: two precise minutes with a mirror beat 15 minutes of vague imitation. Stop once the learner can reproduce the movement in a word. Return after an intervening task and see whether it survives.

Tomorrow's check can be small. Record “The lab analysed 37 samples from men” and “The lab analysed 37 samples from man” with the target noun placed under sentence stress, then ask a listener who has not seen the script to choose the intended sentence and explain which word they heard. Keep the video only long enough to find the movement. Keep the intelligibility test until the end.