Summary
- The Standards and Testing Agency is using an LLM to create draft Year 6 writing for moderator training.
- Human researchers heavily edit every output before internal checks and review by experienced local-authority moderation managers.
- The agency says annual standardisation-production costs have fallen from more than £100,000 to less than £5,000.
The Standards and Testing Agency is using a large language model to produce draft examples of Year 6 pupils’ writing for assessment-moderator training, replacing part of a process that previously cost more than £100,000 a year with one the agency says can be run for less than £5,000.
The Department for Education executive agency has disclosed the system through the government’s Algorithmic Transparency Recording Standard, providing unusually detailed information about where generative AI sits in the workflow, how extensively people rewrite its output, the risks officials have identified, and the savings they expect to make.
Prospective moderators must complete a standardisation exercise before they can approve teacher assessments of pupils’ Key Stage 2 writing, with around 2,000 moderators qualifying each year. The agency previously used a five-year supplier contract to obtain real pupil scripts from schools and assemble collections representing different writing standards.
Under the new process, a large language model is prompted to create material resembling work produced by Year 6 pupils at particular levels within the teacher-assessment framework. Teacher Assessment researchers then heavily edit those outputs before internal review, after which experienced external moderation managers assess the material for authenticity and suitability.
The transparency record says the system is currently being used to create three collections of pupil writing for one standardisation exercise. It identifies the model as “Chat GPT 5 – self hosted” and lists OpenAI as the third party involved, although several procurement, architecture, and data-access fields contain little additional detail.
The economic case is striking because the agency is not claiming a marginal productivity improvement. It says the annual cost of producing standardisation material has fallen from more than £100,000 to below £5,000, while projected production costs for 2026/27 and 2027/28 are about £5,000 a year for quality assurance on two full exercises, one of which will still use scripts retained from the previous supplier arrangement.
Those savings are possible partly because the AI system is not being trusted to produce finished assessment material. Its output is treated as a starting point for specialists who understand the teacher-assessment framework and what real Year 6 writing looks like, while experienced moderators provide another review layer before the samples reach the standardisation process.
The agency has also identified a specific weakness in generated writing. Its transparency record warns that LLM output can omit atypical vocabulary and sentence structures that might appear in work by neurodivergent pupils or children who do not speak English as their first language. Thorough human review and editing are listed as the mitigation rather than an assumption that the model represents the full range of pupil writing accurately.
No personal or sensitive information is supplied to the model, according to the record. Prompts draw on publicly available teacher-assessment frameworks, training material, exemplification resources, and researchers’ knowledge of topics and genres used in Year 6, avoiding the need to expose real pupils’ work to the system as operational data.
The deployment places AI around, rather than directly inside, a consequential decision. The model does not grade pupils, decide whether teachers’ judgements are correct, or approve moderators. It helps create training material that is then rewritten and reviewed by people before being used in an exercise completed by prospective moderators.
That design reduces some risks while leaving others intact. Artificially generated examples still need to be realistic enough that moderators are trained against representative material, and a poorly constructed collection could affect how consistently assessors interpret the national framework even if the model never evaluates an individual child.
The human workload has therefore changed shape rather than disappeared. The old model required sourcing and assembling pupil scripts through an external supplier, while the new process places more emphasis on prompting, specialist editing, internal quality assurance, and external review. Lower procurement cost does not mean the final material can be produced without professional judgement.
The disclosure follows another recent DfE implementation in which AI is being used to check apprenticeship vacancies before selective human review. Together, the records show the department applying automation to narrow operational tasks while retaining human intervention where generated or classified output can affect the quality of a public service.
The transparency record still leaves useful questions unanswered. It does not provide formal model-performance measures, development datasets, an impact assessment, or detailed system architecture, leaving the strongest available evidence operational rather than computational: how people use the model, what officials believe can go wrong, and how much the redesigned process costs.
A decision on whether to continue beyond 2027/28 or return to a supplier model is due in spring 2027. By then, the agency should have several years of training material, quality-review feedback, and actual production costs against which the approach can be judged.












