All articles

Make English intonation visible before copying it

26 September 2026

Choc Education


Two recordings can sound almost identical to the person who made them. On a pitch display, one sentence rises sharply on thirteen and the other peaks on thirty. The line turns a hard-to-hear contrast into something the learner can inspect. For Mandarin-speaking adults refining English intonation, brief visual feedback can make practice more precise.

The screen shows acoustic information. Judgments about whether speech sounds considerate, surprised or clear still come from listeners in context. A teacher and a meaningful sentence have to connect the display to communication.

What a pitch trace shows

Speech software can plot fundamental frequency over time. The trace usually rises as vocal pitch rises and falls as pitch falls. It can help a learner locate the most prominent word, see the shape near the end of a question and compare the timing of a model with a new recording.

Hardison studied computer-assisted prosody training and whether learning transferred beyond practised material. Learners received visual displays with instruction and practice. The results included gains on trained features and evidence of generalisation, alongside qualitative differences in how participants used the feedback. The full article is hosted by Language Learning & Technology, an open journal backed by university language-education centres.

A newer experiment focused directly on English intonation with 49 Chinese university students. It compared visual training, auditory training and a control condition across five sessions in four weeks. The visual group performed better in aspects of intonation form and function. The Lingua paper reports the group sizes and measures.

Forty-nine learners remain a small base for broad prescriptions. The study involved Mandarin speakers at two Beijing universities, so the result is relevant to Taiwanese learners without making the populations interchangeable.

Start with a difference in meaning

A moving line fascinates people for about 43 seconds. Then it becomes another graph. The lesson needs a communicative contrast before the software appears.

Take the sentence “I ordered the BLUE folder.” Prominence on blue can correct a mistaken colour. Move it to ordered, and the speaker may be correcting an assumption that they borrowed the folder. Learners should hear and act on that difference first: choose the right folder, correct a partner or identify which earlier statement the speaker is repairing.

Only then show the pitch traces. Ask where the strongest peak occurs and how long the prominent syllable lasts. The graph now answers a question created by the exchange. Without that exchange, learners can copy a curve beautifully and miss why a speaker used it.

This meaning-first sequence extends the argument in Choc Education's article on how sentence stress guides listeners toward important information. A pitch display supplies feedback on the learner's attempt. The information structure supplies the target.

The same trace can represent several acceptable performances, because speakers distribute pitch across a sentence according to accent, speaking style, emotion and what the listener already knows; the learner's task is to produce a cue that reliably guides this listener in this exchange, then see whether the acoustic display helps explain success or confusion.

Use one feature per round

Pitch displays contain too much data for a beginner. They may show voiceless gaps, tracking errors, tiny fluctuations and a range that differs from the model because of the speaker's physiology. Asking a learner to match the whole shape invites cosmetic imitation.

Choose one feature. In a correction task, inspect the location of the main prominence. In a yes-no question, inspect the final movement only after checking the intended stance. For a list, inspect how non-final items differ from the last item. Leave intensity and segmental detail for another round.

Exact visual matching is a poor goal. A lower-pitched voice and a higher-pitched voice will occupy different parts of the display. Even one speaker changes range with mood and context. Compare relations: where the peak sits, whether the movement rises or falls, and whether the final word carries the intended boundary.

I prefer a rough annotation before a polished graph. Learners can draw a dot above the prominent syllable and a short arrow for the ending. That prediction forces them to listen. Software then checks the prediction with measured speech.

Record, inspect, change

A useful practice cycle takes under four minutes.

  1. Establish the intended meaning with a two-line dialogue.
  2. Listen to one model and mark the expected prominence or final movement.
  3. Record once without watching the screen.
  4. Inspect one selected feature, then record again with one change.

Keep both attempts. Ask a partner to identify the intended correction or attitude from audio alone. If the second pitch line looks closer while the partner understands less, the visual target has taken control of the task and the pair should return to meaning.

Numbers help keep feedback honest. In a set of nine sentences, count how often a listener identifies the intended contrast before and after the learner inspects the trace. A change from four to seven provides information. “The line looks nicer” gives no direction for another attempt.

Software needs supervision

Pitch trackers make errors. Background noise, vocal fry, breathy voice and voiceless consonants can break the line or create octave jumps. A sudden vertical spike may belong to the software. Learners should hear the recording again before trying to repair their voice.

Privacy is another boundary. A web tool may upload recordings to a server. For school use, check retention terms and account settings before asking learners to record identifiable speech. Local software such as Praat can display pitch without turning a classroom exercise into an online archive, though its interface needs preparation.

Visual feedback also suits some targets better than others. It can show pitch movement and timing clearly. It gives limited help with whether a request sounds appropriately polite across contexts, because politeness draws on wording, voice quality, timing and the relationship between speakers.

A 17-minute classroom task

Prepare four short dialogues in which prominence changes the correction: Tuesday versus Thursday, thirteen versus thirty, email versus print, borrowed versus bought. Partners first choose the intended meaning from audio. They mark a peak above one word, record their own line and inspect that word's position on the trace.

After a second recording, partners listen with the screen hidden. They record which meaning they heard. The teacher collects one confusing pair for whole-class analysis and ignores decorative wiggles that carry no communicative load.

End with one transfer line that was absent from practice: “Mina sent the OLD contract.” The learner chooses a correction, predicts the prominence and records once. The useful evidence is the partner's choice between old and another corrected detail. The curve stays on screen for diagnosis, 17 minutes after the first sentence.