Summary
- The Standards and Testing Agency is using an LLM to generate Year 6-style writing that is heavily edited before use in moderator standardisation exercises.
- Its transparency record says annual production costs have fallen from more than £100,000 to below £5,000, although specialist staff and external reviewers remain part of the new process.
- The disclosure gives unusually concrete public-sector AI evidence while also exposing gaps and inconsistencies in the government’s own technical documentation.
The Department for Education’s Standards and Testing Agency is using a large language model to create draft Year 6 writing samples for school-assessment moderator training, with the agency saying annual production costs have fallen from more than £100,000 to below £5,000.
The deployment is documented in an Algorithmic Transparency Recording Standard entry published on 10 August. It describes an LLM being prompted to generate creative writing, factual and fictional prose, and poetry resembling work produced by Year 6 pupils at different points in the Key Stage 2 teacher assessment framework.
The generated text does not move directly into training material. A Teacher Assessment researcher with primary-school assessment expertise heavily edits the output before internal review, after which experienced local-authority moderation managers assess its authenticity. A further quality-assurance process takes place before collections are used in standardisation exercises.
Around 2,000 moderators pass that process each year. Their work supports consistency in teacher assessment of English writing at the end of Key Stage 2, where pupils are judged against a national framework and local-authority moderation is used to check whether schools are applying the standards consistently.
A narrow workflow produces a large saving
The previous production model relied on a five-year supplier contract under which genuine pupil scripts were sourced from schools and assembled into collections representing particular standards. The agency says that process cost more than £100,000 annually, created a burden around acquiring scripts, and could still produce examples that sat awkwardly between assessment boundaries.
Generative AI gives researchers a different starting point because prompts can combine the assessment framework with topics and genres familiar to Year 6 pupils. The output is then treated as raw material rather than a finished pupil simulation, allowing researchers to shape it before experienced moderators determine whether the resulting work appears authentic.
STA says the LLM is currently being used to produce three collections for one standardisation exercise. For 2026/27 and 2027/28, it expects annual production costs of around £5,000 for quality assurance on two full exercises, one of which will continue to contain scripts produced through the previous supplier arrangement.
The claimed reduction is unusually specific for a government generative-AI deployment, although it should not be read as replacing more than £100,000 of work with a £5,000 software bill. Internal researchers still carry out extensive editing, external moderation managers provide specialist review, and the transparency record does not calculate the full staff cost attached to the new workflow.
Human judgement carries the assurance burden
The agency acknowledges that generated writing may fail to represent the range of language found among real pupils. Its risk section specifically notes that outputs could omit atypical vocabulary and sentence structures associated with neurodivergent children or pupils for whom English is not a first language, with thorough human review identified as the mitigation.
That makes editing a substantive control rather than cosmetic proofreading. If synthetic examples converge on a narrower idea of what Year 6 writing looks like, moderators could be trained on collections that appear plausible while representing less variation than the classrooms they eventually assess.
No personal or sensitive pupil data is supplied to the model, according to the record, while prompts draw on the teacher assessment framework and researchers’ knowledge of age-appropriate topics and genres. The process therefore avoids using children’s personal information as model input, even though its purpose is to generate material resembling children’s work.
The transparency record is also revealing for what it fails to resolve. It identifies “Chat GPT 5 – self hosted” as the model, lists OpenAI in another field, records third-party involvement as “No”, and leaves architecture, datasets, maintenance, procurement, and several other technical fields blank. The operational workflow is consequently documented more clearly than the technical system underneath it.
That inconsistency does not mean the deployment is necessarily unsafe, but it weakens the transparency standard’s ability to answer basic questions about how government AI is being supplied and operated. A public record is considerably more useful when departments describe the same system consistently across ownership, procurement, hosting, and model fields.
The use case also sits beside other bounded applications inside the department. DfE has disclosed AI use in apprenticeship vacancy quality checks, where automation similarly sits ahead of selective human review rather than replacing the final administrative process.
Moderator training offers a stronger case study than broad promises to automate education administration because the workflow is narrow, the old process is known, and the agency has attached an explicit cost comparison to the change. Yet the same disclosure also shows why human expertise remains central: the LLM produces material cheaply, while experienced people still decide whether it resembles the pupils whose writing the assessment system is supposed to understand.












