91av

Test moderators use AI-generated writing to judge literacy standards

To cut the cost of key stage 2 assessments, a UK government project proposes to use scripts generated by ChatGPT instead of real children's writing as examples supplied to moderators who check that grades are awarded fairly
Key stage 2 tests are taken by pupils in year 6
Catherine Delahaye/Getty Images

The UK Department for Education (DfE) is testing the use of AI to produce writing samples used in the assessment of pupils’ literacy standards as they move between primary and secondary school, claiming it cuts annual costs by 95 per cent.

Around 2000 moderators are asked to check the standard of pupils’ writing at the end of key stage 2 (KS2), a critical moment in a child’s educational pathway in England and Wales. To ensure that grades given by teachers are standardised, the moderators cross-check pupils’ work against samples of writing usually drawn from actual work produced by schoolchildren. That process costs around £100,000 a year.

Starting this year and next year, some moderators will AI-generated output created by ChatGPT’s GPT-5 model.

The experiment is limited for now. DfE says GPT-5 is currently being used to create three collections of writing for one standardisation exercise, while one full exercise in 2026-27 and 2027-28 will continue to contain scripts written by children and obtained under the previous contract.

Around 20 experienced local authority moderation managers will review the AI-derived material for authenticity before it is used. The department plans to decide in spring 2027 whether to continue producing samples this way or return to procuring them from an external supplier. Its transparency record lists no formal impact assessment for the system.

However, the use of AI in a key part of educational assessment has raised questions.

“Having an exemplification of writing that is not a real child’s writing for me creates a philosophical and ethical issue,” says at Anglia Ruskin University, UK. “The worry is that AI shifts our perspective of what’s acceptable in writing, or what is a good standard, or something like that.”

Clarkson, who studies KS2 writing assessment and moderation, has heard from a number of moderators who have raised concerns about the potential use of AI-generated samples. “Essentially [their criticism] just centres around the fact that it’s not real children,” she says.

The consequences of this are impossible to predict, she says. “Is it a concern if we don’t use real children’s writing? Who knows? I guess we don’t know what the impact is until we can measure the impact.”

The DfE didn’t respond to 91av’s request for comment.

The use of AI in this way shouldn’t affect which secondary schools pupils end up at, since the KS2 writing assessments take place after pupils have been assigned to their new schools. But how well children match up to the synthetic samples could dictate the level of support they receive when they enter their next school.

Because AI can tend to produce flat or bland writing, there are worries that children who speak English as a second language or who are neurodivergent and so may not use the same language as the majority may be discounted. DfE’s notes that “LLMs produce texts that often [exclude] atypical vocabulary and sentence structures that might be used by neurodivergent or non-native [English-speaking] students” – something it plans to mitigate by “thorough review”.

The assessment also points out that the AI-generated outputs will be “heavily reviewed and edited” and subject to review before being used in standardisation processes.

“Use of AI is seeping into all aspects of our lives and without proper governance, we do not know what the implications will be,” says at the University of Oxford. “In assessment, there is a concern that there will be a backwash, with the AI-generated materials becoming the standards that teachers try to get pupils to emulate. This could lead us into some strange places.”

Topics: AI / education