Summary
- JRC prototypes used generative AI to extract and connect information from outbreak reports, news sources, and epidemiological records.
- The systems could synthesise unstructured evidence faster, but they have not been validated in operational public-health surveillance.
- Researchers favour phased pilots backed by common data standards, bias controls, validation, and continued human responsibility.
European researchers are testing whether generative AI can shorten the laborious process of piecing together early signs of disease outbreaks, although the work remains some distance from putting large language models into operational public-health surveillance. The European Commission’s Joint Research Centre has developed prototypes that process large volumes of fragmented epidemiological information and assemble it into more usable intelligence for specialists monitoring emerging health threats.
The research addresses a problem that predates the current enthusiasm for generative AI. Useful warning signals can be scattered across official outbreak reports, news coverage, specialist databases, and other unstructured sources long before they appear in a consistent dataset, leaving epidemiologists to connect evidence produced in different formats and languages. The JRC found that generative systems could accelerate some of that synthesis, potentially reducing the time spent organising information before experts decide whether a signal warrants investigation.
Researchers worked with material from the Epidemic Intelligence from Open Sources system, known as EIOS, which aggregates publicly available information used in disease surveillance. One part of the work created a knowledge graph linking epidemiological information across sources, while another used retrieval-augmented generation with World Health Organization Disease Outbreak News to produce richer accounts of emerging health threats.
However, the JRC is explicit about the limits of the findings. The prototypes have not been validated in operational public-health environments, which means researchers do not yet know whether improvements demonstrated during development will translate into faster or more dependable detection once systems encounter incomplete evidence, multilingual information, organisational constraints, and decisions carrying substantial public consequences.
Speed does not settle the clinical question
Those limitations are particularly important in epidemiology because a plausible but incorrect AI-generated interpretation can create work rather than remove it. Large language models can produce confident answers from incomplete or contradictory evidence, while public-health teams need to establish why a signal was elevated, what information supports it, and whether a suggested connection between events can withstand expert scrutiny.
The JRC consequently treats generative AI as an analytical aid rather than an autonomous detection system. Human specialists would continue to determine whether emerging signals are credible, what additional evidence is required, and what response should follow, while software could absorb part of the information-management workload upstream of those judgements.
That model resembles the way generative AI is beginning to enter other regulated environments, where bounded automation can be easier to govern than handing an entire decision chain to a model. Public-health surveillance provides a particularly demanding test because false negatives can allow a threat to develop unnoticed, while false positives can divert laboratories, staff, and public resources towards events that do not justify intervention.
Governance therefore becomes part of the system design rather than an administrative layer applied after deployment. The JRC identifies validation, bias mitigation, privacy safeguards, and interoperable data standards among the conditions for wider use, while expanded digital surveillance also raises questions about how information is collected and combined across national and institutional boundaries.
Fragmented data remains the harder constraint
Even a highly capable model cannot remove the fragmentation of European public-health systems. Member States collect, structure, and exchange information differently, so common formats and interoperability become part of the infrastructure required for useful AI-assisted surveillance. Without them, a new model risks becoming a sophisticated interface placed over inconsistent inputs.
Retrieval-augmented generation offers one way of narrowing that problem. Rather than asking a general-purpose model to answer only from information embedded during training, a RAG system retrieves material from defined sources before generating its response. That gives public bodies greater control over the evidence presented to the model and can make it easier for specialists to trace an output back to an underlying record, although it does not eliminate hallucination or interpretation risk.
The JRC is therefore proposing a phased route towards adoption rather than an EU-wide deployment. National or regional agencies could run pilots in real surveillance environments, allowing teams to measure whether the systems improve detection and response, how often specialists need to correct them, and whether any productivity gain survives contact with multilingual and unevenly structured data.
Successful trials could then inform wider adoption supported by shared standards, governance, and training. That is a more demanding proposition than buying a generic AI service and connecting it to an existing workflow, since agencies would need evaluation methods covering not only model accuracy but also escalation procedures, data protection, staff workload, and the time required to verify machine-generated conclusions.
The work has involved public-health stakeholders including the European Centre for Disease Prevention and Control, the World Health Organization, the Commission’s health directorate, and the Health Emergency Preparedness and Response Authority. Their involvement gives the research a route towards operational testing, although the current findings remain evidence about prototypes rather than proof that generative AI is ready to become part of Europe’s routine epidemic-intelligence machinery.
Operational pilots will provide the harder evidence. Processing thousands of documents more quickly is technically useful, but public-health agencies will ultimately judge the systems by whether they help specialists make dependable decisions while preserving the ability to challenge an output, inspect the evidence beneath it, and remain accountable for the action that follows.












