Written English can steady a new word in memory. It can also pull the pronunciation toward the wrong sound. A useful vocabulary lesson controls the order: learners hear a new word clearly, attempt to identify its sounds, then see the spelling and connect both forms to meaning. Showing everything at once gives the page too much authority, especially for Taiwanese learners whose English education has often rewarded accurate spelling more visibly than accurate listening. A brief delay lets the sound establish its own trace before letters enter the picture.
The page should wait.
Spelling leaves a strong trace
Speech disappears as soon as it is heard. A printed form stays available for inspection. Learners can count letters, notice a repeated ending, and compare the word with something already known. That stable record can support memory. Salomé, Commissaire, and Casalis tested this with French children learning German words. In one experiment children learned 16 items; in another they learned 24. The 2024 study included third- and fifth-grade developing readers and compared words learned with orthography against words learned without it. Seeing the written form helped on written choice, spoken recognition, and the connection from sound to meaning. Both age groups benefited.
These were children learning small sets of German words in a paired-associate experiment. The result cannot set a universal procedure for Taiwanese adults learning English through conversation. It does show that print may support spoken-word learning rather than belonging to a separate reading box. Suppose an A2 learner hears aisle while preparing for a trip. Audio plus a supermarket picture establishes a first sound-meaning link. The spelling then gives the learner a stable label to revisit. If the lesson begins with A-I-S-L-E, however, the silent s may receive a sound that the recording never contained.
Order matters.
The same anchor can pull in the wrong direction
English spelling offers unreliable clues for many vowels and consonants. Steak, break, and speak look closely related. Two rhyme; one does not. Receipt carries a silent p. Women starts with the spelling of woman while its first vowel sounds different. A learner who meets these words first on a page may build a precise memory for an inaccurate target. Pauline Welby, Elsa Spinelli, and Audrey Bürki followed French-speaking adults across two days as they learned novel English words. Half appeared with spelling and half through audio alone; learners also heard either one speaker or several. On day three, spelling produced faster responses in naming and recognition. Acoustic measurements revealed the cost: written forms pulled produced vowels toward French sound categories. Pronunciations became more consistent, yet their position in the vowel space was less English-like.
The items were invented, the first language was French, and the study examined a controlled learning period. Mandarin speakers bring different sound and writing systems. The finding still captures a useful tension. Orthography can strengthen access to a word while biasing the sound representation that access retrieves. Choc's article Let sound lead when English spelling lies explains the problem for established words. New-word teaching can prevent part of it by giving sound a brief lead before print arrives.
A thirty-second sound-first sequence
Take colleague, a word that may appear in workplace material. Keep the written form covered at first. Play or say the word twice in a short message: My colleague in Tainan will send the figures. Ask learners to choose the person in a simple picture or identify the relationship. Play it again and let them mark two beats with their fingers: COL-league. Now reveal colleague. Learners compare what they expected with the page, underline the stressed first syllable, and cross out no letters because every letter still belongs to the conventional spelling even when the vowel-letter relationship surprises them. They say the whole message, then retrieve the word from the picture after 11 minutes.
The sequence has four jobs:
- audio establishes a sound target;
- context establishes meaning;
- print makes the form available for study;
- delayed recall checks whether meaning can activate the spoken word.
Twenty or 30 seconds of sound-first work is enough for one item. Turning every vocabulary lesson into phonetic analysis would consume the message the word was chosen to express. Words with transparent spelling, such as fantastic, need less protection. Words with a likely trap deserve more: comfortable, choir, debt, island, recipe, vehicle.
Keep spelling and sound connected afterward
Delaying print once does little if every later review is silent. A digital card can present a picture or meaning cue, require a spoken answer, play a recording, and show spelling last. A paper card can place a small stress mark beside the word and a short phrase underneath. The learner should hear and say a comfortable chair, since the rhythm of the phrase may reduce the word more than a careful dictionary recording does. Assessment should separate access from accuracy. One check asks the learner to point to the right picture after hearing the word. Another asks for the word after seeing the picture. A third shows the spelling and asks for pronunciation inside a phrase. If picture-to-speech succeeds while print-to-speech produces an extra consonant, the learner knows the meaning and needs a repair between letters and sound.
Teachers also need to respect learners who depend on text for accessibility or memory. Removing spelling for an entire lesson can raise load and frustration. I prefer a short delay measured in seconds, followed by full access to the word. The procedure changes when a learner has hearing loss, when the target is technical spelling, or when accurate written production is the immediate goal.
Seconds are enough.
Choose six words from the next listening text and predict which spellings could mislead. Let learners hear those six in meaningful phrases before opening the transcript. Reveal the page, compare expectation with print, and return to audio once more. Spelling then serves memory without getting the first vote on pronunciation.