Summary
- Regnology surveyed 276 practitioners across 22 countries and found widespread AI exploration and pilot activity.
- Only 16% described AI as embedded operationally, with large institutions showing little advantage in production adoption.
- Explainability, auditability, data quality, governance, and supervisory expectations remain barriers between experimentation and trusted use.
Financial institutions are finding it considerably easier to pilot artificial intelligence in regulatory reporting than to trust it with the reporting cycle itself, according to new research from Regnology.
The regulatory technology company surveyed 276 practitioners across 22 countries and found that 87% were exploring, piloting, or embedding AI in their operations, while only 16% had reached embedded production use. The gap was not confined to smaller institutions with fewer technical resources.
Across institution sizes, embedded use remained within a narrow range, suggesting that additional money and technical staff make experimentation easier without necessarily solving the governance questions involved in production deployment.
Regulatory reporting creates a particularly demanding test because automation has to coexist with auditability, accountability, and the possibility that supervisors will ask an institution to explain precisely how a number was produced. A system generating plausible output is insufficient when a bank must reconstruct the data lineage, controls, and decisions behind a filing.
Modernisation comes before autonomy
Regulatory-reporting operations commonly combine legacy systems, manual reconciliations, spreadsheet work, repeated data transformations, and specialist interpretation. Adding an autonomous agent to that environment does not remove the inconsistencies underneath it; without strong traceability, automation can make them harder to see.
The production question therefore extends beyond whether a model can perform the task. An AI tool summarising an exception for a human analyst creates a different risk profile from an agent that changes data, resolves the exception, and moves the result into a controlled reporting workflow.
Regnology describes the distance between a pilot and trusted operational use as the “agentic gap”, arguing that explainability, auditability, domain translation, supervisory expectations, data, and governance all have to be addressed before autonomy can increase safely.
Those constraints help explain why the largest institutions can lead on pilot activity without pulling dramatically ahead in embedded use. Their scale provides data, engineers, and experimentation budgets, but it also brings larger technology estates, more controls, and more reporting obligations.
The economics remain conditional
There is a clear reason banks continue to experiment. Regulatory reporting absorbs substantial operating expenditure, and many processes contain repetitive work that appears technically suitable for automation.
Regnology’s research, supported by commissioned Oliver Wyman analysis, suggests a meaningful share of reporting expenditure may be addressable through agentic workflows. That is a directional estimate rather than a guaranteed saving, and expenditure theoretically touched by automation is not the same as cost that disappears from an institution’s accounts.
Regulated workflows often require parallel controls, validation, and human oversight after automation is introduced. During migration, expenditure can rise because institutions are paying for new systems while maintaining older processes, while stronger model governance creates work around documentation, testing, and monitoring.
The more practical deployment question is where automation can remove repeated effort without weakening the control environment. Reconciliation, anomaly triage, and initial analysis may lend themselves to progressively greater automation, while material judgements and final attestations are likely to retain human accountability for longer.
Trust becomes an engineering requirement
Financial services provides an early indication of the difficulties other regulated sectors will encounter as agentic systems move closer to production. Healthcare, utilities, government, and critical infrastructure all contain workflows where model accuracy is only one part of the deployment decision.
Organisations also need to know what data an agent used, what permissions it held, what actions it took, and who remained accountable when the process moved outside expectations. Those requirements push AI projects towards identity management, logging, access controls, testing, and process design rather than model choice alone.
A bank may be able to build a technically capable agent in weeks while taking much longer to integrate it into a governed reporting process spanning several jurisdictions and regulatory regimes. The difference between those timelines explains why enterprise AI adoption can look rapid when measured through pilots and much slower when measured through production authority.
Financial institutions are plainly willing to test AI in high-value reporting work, but production deployment still requires them to decide where software can act independently, where it should remain supervised, and how every consequential step can be reconstructed afterwards.
Until those conditions become routine engineering rather than bespoke governance work, pilot counts will continue to overstate how deeply agentic AI has entered regulated operations.












