すべての記事

Transcribe one recording before you speak again

2026年9月29日

Choc Education


An English recording becomes far more useful after the learner writes down exactly what it contains. For a 34-year-old project manager in Neihu who can prepare a slide deck yet loses control of verb endings while presenting it, one short self-transcription can expose problems that another round of free speaking leaves hidden. The best next step is to record a manageable task, transcribe the actual words, mark two repairable gaps, and perform the same task again.

Speech disappears too quickly

During live conversation, attention has several jobs. A speaker has to choose an idea, retrieve words, arrange them, pronounce them, watch the listener, and decide what comes next. A missing article or a muddled tense may pass before the speaker can inspect it. The recording slows that moment down. Transcription is deliberately literal. If the recording says, “Yesterday we discuss the revised timeline,” the page should say discuss, without silently correcting it to discussed. That small discipline matters because the visible gap comes from the learner's own production. A grammar worksheet may contain the same past-tense form, though it cannot show where that learner dropped it while explaining Milestone 7 to a colleague.

This is inspection, not punishment.

What the studies found

Tony Lynch compared two transcript procedures in two English for Academic Purposes classes. Learners completed a role-play, then one class transcribed its own performance while another worked from extracts selected and transcribed by the teacher. They performed the task again two days later and once more four weeks later. Lynch's 2007 study found that both procedures could be run under normal classroom conditions; the self-transcribing class appeared better able to maintain accuracy in the forms highlighted during the follow-up work. The design does not prove that transcription improves every part of speaking. It compared two classes, used a repeated role-play, and concentrated on the forms learners had noticed. Fluency, spontaneous transfer to a new topic, and pronunciation need separate evidence.

A later classroom study gives a second, longer view. In Jafari and colleagues' 2016 study, intermediate and advanced adult EFL classes recorded group discussions over 20 weeks. Experimental classes transcribed the conversations, first corrected individually, and then discussed and reformulated inaccurate utterances with peers. Control classes recorded the same kind of discussions without that post-task work. The transcription-plus-correction classes improved their grammatical accuracy. Several parts of the treatment moved together, so the result cannot be assigned to transcription alone.

That caveat changes the classroom recipe. A transcript left in a folder is paperwork. The productive sequence includes noticing, checking, correction, and another attempt.

The boundary matters.

Keep the sample short

A three-minute recording can take much longer than three minutes to transcribe. Lynch reported detailed learner work that demanded real time, and anyone who has paused a voice message after every five words knows why. Asking a busy adult to transcribe a full 18-minute mock presentation will probably teach fatigue. Start with 45 to 75 seconds. Choose a task with a stable purpose: explain one chart, leave a voicemail changing an appointment, or summarise the decision from a meeting. Automatic transcription may provide a rough first pass, but the learner should listen against it. Speech recognition often regularises weak grammar, misses filled pauses, and guesses the wrong word precisely where the recording deserves attention.

Use three passes:

  1. Write the words that were actually spoken, including repetitions and unfinished starts.
  2. Underline two places where meaning, grammar, or pronunciation broke down. Check a reliable reference or ask a teacher before changing uncertain cases.
  3. Prepare a corrected version, put it away, and record the same communicative task again.

Two targets are enough. A page covered in twelve correction symbols makes selection impossible, especially for an A2 speaker still spending most of their attention on retrieving basic words.

Listen for meaning as well as grammar

Learners often begin by hunting visible grammar errors. The recording may reveal a larger problem. “We can deliver Friday” could be accurate English and still leave the listener unsure whether Friday is a promise, a proposal, or the earliest possible date. A transcript permits a second question: did each sentence do the intended job? Marking can separate three kinds of repair. Use G for a form that needs checking, W for a word or phrase that obscures the idea, and L for a listening problem such as a swallowed ending or misplaced pause. This is a preference, not a standard notation system. It keeps the page usable.

Meaning comes first.

Pronunciation comments need the audio beside the text. A learner may write the correct word worked while producing an ending that disappears. Conversely, a non-native accent can remain fully comprehensible and require no repair. Correct the feature that cost the listener meaning.

Repeat the communicative task

The second recording should preserve the purpose without becoming a memorised recital. Keep the same chart but cover the corrected script. Keep the same voicemail problem but change Tuesday to Thursday and 3:30 to 4:15. Those small changes require the learner to rebuild the message while giving the repaired language another chance to appear. This connects with Choc's article on speaking-task repetition. Repetition reduces some planning pressure. Self-transcription adds a focused inspection between attempts. Together they give the learner a reason to repeat beyond trying to sound vaguely smoother. Recorded monologues, paired role-plays, and rehearsed workplace explanations fit the method. A fast, emotionally loaded conversation fits poorly because recording may alter the interaction and consent matters whenever another person is audible. It also asks too much of a beginner who cannot yet segment their own speech into words. A teacher transcript containing 12 selected seconds may be a better starting point there.

Progress can be logged without grading the speaker's personality. Keep both recordings, note the two chosen targets, and check whether those forms survive in a fresh 70-second task the following week.

Tomorrow morning, record a 60-second explanation of one chart. Transcribe every word before opening the slide deck again. Circle two places, check them, then explain the same chart with the transcript face down.