All articles

English self-assessment should name the task

15 September 2026

Choc Education


English self-assessment should name the task. "My English is intermediate" gives a learner a broad location. It says little about whether she can follow a weather report, explain a production delay on Microsoft Teams, or write a 120-word reply to a customer. A useful self-assessment breaks proficiency into actions that can be attempted, recorded, and checked against a sample of performance, then uses the mismatch to choose one narrower piece of work for the next study session. This article is for a 31-year-old office worker in Neihu who calls herself B1, studies on the MRT, and needs clearer evidence for choosing next month's work. A level label can stay. The study plan should begin one layer below it.

Levels compress uneven skills

Language ability has a jagged profile. The same learner may follow familiar project updates, lose a fast exchange between two colleagues, write accurate routine email, and avoid an unscripted phone call. One global label averages those differences. It helps with course placement and communication between institutions. It gives a weak instruction for Tuesday's 27-minute study session. A task statement narrows the decision: "I can identify the time, place, and required action in a 45-second voicemail after two listens." The learner can select material, try the task, and keep the result. If the place is correct and the required action is missing, tomorrow's work has a target. Confidence alone cannot supply that detail.

"Improve listening" is still too large.

Taiwan has unusually useful evidence

The Language Training and Testing Center developed GEPT self-assessment statements for Taiwanese learners and tested how well the judgments matched test performance. Jessica Wu and Chia-lung Lee report a study inviting 8,006 Taiwanese EFL learners to take a GEPT test and answer 22 listening and 21 reading can-do statements. The reported accuracy for estimating level was 0.68 in listening and 0.65 in reading. Those figures are encouraging and sobering. A short set of concrete statements predicted broad test level far better than a casual guess would, yet prediction remained imperfect. Self-assessment can help a learner choose a likely GEPT level and notice a weak area. It should not replace an official result where a school, employer, or licence requires one.

LTTC's public tool makes the wording visible. Its GEPT self-diagnosis page says the statements were built from GEPT ability descriptors and empirical data. The listening items ask about actions such as understanding prices while shopping, extracting times and locations, following simple instructions, or inferring a speaker's attitude. These are more useful than "good at listening" because each item points towards a class of audio and a possible check. The limits deserve attention. The research linked self-reports with a standardised test, so alignment partly depends on how closely the statements represent the GEPT construct. A learner may handle a familiar workplace call well and still misjudge an unfamiliar news item. Self-assessment also shifts with wording, recent success, anxiety, and knowledge of the task.

Calibration takes evidence.

Practice changes the assessor

Yuko Butler and Jiyoon Lee studied 254 sixth-grade learners of English in South Korea. Students used self-assessment regularly during one semester. Their ability to assess their performance improved over time, and the study found small positive effects on English performance and confidence. Teachers' views of assessment shaped how they perceived and implemented the practice. The participants were children in South Korea, so the size and form of the effect cannot be carried straight into an adult Taiwanese course. The more defensible lesson concerns repetition: self-assessment is a skill. Learners need chances to predict, perform, compare, and adjust. A single end-of-term confidence survey offers no such calibration. I would rather see four modest judgments with evidence than one dramatic claim about fluency. The evidence can be plain: a timestamp where the message became unclear, a recording, two missed details, a teacher rating, or the difference between a predicted and actual quiz score. The purpose is to improve the learner's next judgment.

Write statements that can be tested

Good task statements include four pieces: an action, content, conditions, and a success threshold. Compare "I can speak about work" with "I can give a 90-second update on a familiar project, explain one delay, and answer two follow-up questions without notes." The second statement can produce a recording and two questions. It also reveals which part failed. Avoid packing unrelated skills into one line. "I can understand a podcast and discuss it accurately with natural pronunciation" combines listening, recall, interaction, grammar, and speech. A "no" answer gives no diagnosis. Split the work according to the decision the learner needs to make.

A compact weekly set might contain four statements:

  • I can catch the gate, boarding time, and delay reason in one airport announcement.
  • I can leave a 35-second message that includes my name, problem, and requested reply.
  • I can read a 240-word product notice and find two conditions attached to the refund.
  • I can write a six-sentence update using past time consistently.

Those numbers create stable conditions. They do not define language ability for all time. A 35-second message can be repeated next week with different content, which makes comparison possible without turning practice into the same memorised script. The learner can also keep the format while raising one demand, such as removing notes, adding a follow-up question, or reducing the second listening. Change one condition at a time and the source of improvement stays visible.

Pair judgment with performance

Before the task, record a prediction on a simple scale: "easy," "possible with effort," or "not yet." Perform the task. Check the result against a transcript, answer set, recording, peer response, or teacher comment. Then write one sentence explaining the mismatch. This short cycle separates confidence from evidence without treating confidence as a defect. The Council of Europe's 2020 CEFR Companion Volume organises descriptors across specific communicative activities, including conversation between other people, announcements and instructions, audio media, correspondence, online interaction, and mediation. It also warns against reducing the CEFR to a gatekeeping instrument. The descriptor is a reference point for learning and assessment, while the learner remains a person acting in a social setting. Low-stakes quizzes can add another comparison point. Frequent testing can support learning when the stakes stay low, and the prediction beside each result adds metacognitive evidence: "I expected 8/10 and scored 5/10 because I missed reduced forms." The score alone would hide the mismatch.

The comparison belongs after performance.

Know when a test is needed

Self-assessment suits planning, reflection, and the choice of practice material. A formal decision may require a validated assessment under standard conditions. University admission, a promotion rule, or a GEPT certificate carries consequences that a personal checklist cannot bear. There is also a case for skipping self-ratings during fluent activity. A learner who judges every sentence while speaking may slow down and avoid risk. Record the task first. Review it later. Reflection belongs after some performances rather than inside every turn. A weekly review is enough for many learners, and some weeks may yield no useful judgment at all.

Replace one label this week

Keep "B1" at the top of the page if it helps. Under it, write one task needed before Friday: "I can explain one production delay in 90 seconds and answer two follow-up questions." Predict the result, make the recording, and listen once with the four required parts on paper. The next study session can begin at the missing part rather than at the word intermediate.