Summary
- DSIT used its Consult AI system to identify themes in open-text responses to a national consultation.
- Civil servants reviewed, edited, and rejected machine-generated themes before using them in the analysis.
- Detailed email submissions were reviewed separately, showing automation being inserted into part rather than all of the evidence process.
Artificial intelligence has entered a less visible part of British policymaking, with civil servants using a government-built system to help turn thousands of free-text consultation responses into themes that officials can examine before ministers make policy decisions.
The Department for Science, Innovation and Technology has disclosed how its Consult tool was used during analysis of the “Growing up in the online world” consultation, an engagement exercise that received 116,211 responses across questionnaires, surveys, campaign submissions, and other channels.
Consult does not simply produce a summary that officials accept. After responses are prepared for analysis, the system identifies potential common themes, which civil servants compare against the underlying submissions. Officials can edit themes they consider inaccurate, reject those regarded as irrelevant or repetitive, and then use the tool to count references to the agreed categories.
The deployment puts AI inside a long-established but labour-intensive administrative process. Consultation analysis traditionally requires teams to read and classify large bodies of written material before policy officials can compare recurring arguments and evidence.
Automation changes the middle of policymaking
Government AI projects are often discussed through citizen-facing chatbots or internal productivity tools that summarise meetings and documents. Consultation analysis sits in a more consequential middle layer because the resulting themes become part of the evidence officials use when advising ministers about public and organisational responses to proposed policy.
That makes the structure of human oversight more important than a generic claim that a person remains “in the loop”. In this case, officials review themes generated by the system against source material, modify or reject them, and retain responsibility for interpreting the resulting evidence.
DSIT also altered its standard process because of the scale and timing of the exercise. An interim analysis using roughly the first 12,000 responses established initial themes, which were then applied during work on the wider dataset. Officials used additional checks intended to identify later responses containing new arguments or particularly detailed evidence.
The approach can reduce the cost of repeatedly rebuilding the thematic structure, although it introduces a potential analytical risk: categories established from earlier responses may influence how later material is interpreted. The department describes additional flags, human review, and consistency checks intended to mitigate that effect.
Not every submission went through the model
The published methodology provides useful boundaries around the automation. The consultation received 279 unique email responses, which officials reviewed individually because they were more likely to contain detailed evidence and did not follow the questionnaire structure.
A further 33,141 campaign emails formed another distinct part of the dataset, while separate questionnaires and panel surveys gathered views from parents, children, and young people. The headline figure of 116,211 responses therefore should not be read as meaning that Consult independently analysed 116,211 equivalent free-text documents in a single process.
The technology is being used to accelerate theme identification and counting within a broader analytical workflow, rather than replacing consultation analysis or deciding what policy should follow from the submissions.
The evidence has limitations regardless of the technology. DSIT explicitly notes that consultation respondents are self-selecting and should not be treated as a representative opinion poll. AI can organise the material more quickly but cannot make an unrepresentative sample representative.
Transparency becomes part of AI assurance
Using AI in a process feeding into public policy also creates pressure for government to explain the system with enough detail for outsiders to understand what was automated. Publishing the methodology reveals the sequence of preparation, theme generation, official validation, counting, and additional manual review.
That procedural information is more useful than simply stating that AI was used with human oversight because the principal risks sit inside the hand-offs. A system could overlook minority viewpoints, merge distinct arguments, over-count similar wording, or struggle with unusual submissions, while reviewers might reinforce those errors if generated categories are treated as more authoritative than the source material.
Having officials validate proposed themes and separately inspect evidence-rich responses addresses some of those risks, although confidence will depend on whether similar methods remain consistent when Consult is used by other departments and on different kinds of policy question.
The deployment nevertheless represents a more concrete use of government AI than many trial announcements because it has entered a live administrative workflow producing published policy evidence. The potential gain is straightforward: fewer staff hours spent manually classifying large volumes of text and more time available to examine what the responses actually contain.
That saving remains useful only when traceability, minority arguments, human judgement, and disclosure survive the automation. Consult can make the sorting faster, but the legitimacy of the resulting evidence still depends on the process surrounding the model.












