
The central issue in the debate over introducing essay-based questions to the College Scholastic Ability Test, or CSAT, now being pursued by the presidential National Education Commission, is grading fairness. In a college admissions system where a single point can decide who gets in, critics warn that fairness will be questioned unless students can be convinced why points were deducted. Some have gone so far as to argue that discrepancies between graders must be driven to zero.
Grading fairness must of course be secured. But reducing error to zero would require an exam with a single fixed answer, which is no different from a multiple-choice test. Consider how other major countries grade essay-based college entrance exams. Britain's A-levels, France's baccalaureate, China's gaokao and the International Baccalaureate all operate standardized essay-grading systems. None of them assumes that an essay has one correct score. Rather than making a one-point difference explainable, they build systems that deliver consistent scores within an allowable range no matter who does the grading.
IB examiners, for instance, must pass a grading test for every exam session, meaning that even someone who served as an examiner last year must pass again this year. They practice on benchmark scripts to calibrate their standards. Benchmark scripts are randomly inserted into each set of 10 answer papers, and if an error is detected the examiner is automatically logged out and the set is reassigned to another examiner. Disagreements in cross-grading are settled by a senior examiner, and students who dispute their results can request a re-mark. These layered mechanisms bring grading around the world into line with a single standard. China's gaokao also averages the scores of two graders, and if the gap between them exceeds an allowable range, a third grader and experts re-mark the paper. This year 12.9 million students sat the exam, and trained teachers and professors finished grading within two weeks. There is no reason Korea, with 500,000 test takers, cannot do the same.
Easing the sensitivity of a single point also requires broadening the scope of assessment. The Korean language section of the CSAT ends with a single sitting, while IB language courses combine passage analysis, comparative essays, essays and oral assessment. External assessments are graded blind, while internal assessments graded by teachers are randomly sampled and reviewed and adjusted from outside. Repeated verification across multiple points in time, formats and assessors raises the stability of measurement and disperses the weight of a single point, so that outcomes are not excessively determined by one exam or one person's judgment. Raw scores summing results from various angles are published alongside grade scores, but the system has earned enough credibility that a considerable number of universities worldwide use only the grade scores, which have passed through multiple layers of cross-verification, rather than the finer raw scores.
Consistency alone, of course, does not complete fairness. The grading criteria themselves must be valid. The writing criteria for China's gaokao include creativity but also explicitly list soundness of thought, so no matter how creative a piece is, it cannot score highly if its ideas are judged unsound. Britain's A-levels are under pressure to reform amid criticism that they favor breadth of knowledge over thinking skills. Comparative studies show that even among essay-based formats, how deeply thinking skills are measured differs markedly from one assessment system to another. Determining what to measure is as decisive a factor in assessment fairness as consistent grading.
There is no perfectly fair assessment in the world, only efforts to design a fairer system. Grading is a kind of rule of the game. Double the time allowed on the CSAT and the rankings change. Because the CSAT measures not only what students know but also their speed, it gives the same score to a student who could not write an answer out of ignorance and one who ran out of time. That is not fair, yet it is accepted because the same rules apply to everyone. The same holds for essay grading. If Korea benchmarks proven systems and builds consistent rules through the standardizing steps of benchmark scripts and error correction, whether by humans or by artificial intelligence, the credibility of assessment can be sufficiently secured.
Fairness in assessment does not come from eliminating the grader's judgment. It comes from designing systems in which that judgment is valid and consistently reproduced. The essential choice before the National Education Commission is not between zero error and some error. It is between an exam that hides error and an exam that admits error and manages it precisely. The essay-based college admissions systems of major countries have already chosen the latter.






