Builder's Notes · · 7 min read
When AI Meets Regulatory Data: Three Correct Answers to One Simple Question
The difficult part of regulatory intelligence is rarely the model. It is defining exactly what every number means before the model is allowed to answer.
By Jan Kluz, Founder, Emable
Jan builds data and AI products for European financial distribution, with a focus on advisor intelligence, decision validation and responsible use of regulatory data.
Ask a financial-market database how many advisors exist and it can return several different numbers without making a single technical error.
That sounds impossible until you look at what a regulatory register actually records. The Czech National Bank register is designed to document authorisations. One person may hold several authorisations, across several product categories and sometimes through more than one company. The register is authoritative for its purpose, but its natural unit is not always the human being a manager has in mind.
This distinction is the starting point for responsible AI on regulatory data.
A correct number can still answer the wrong question
When someone asks, “How many advisors are in the market?”, they may mean at least three things:
- How many authorisation records are currently visible?
- How many distinct people appear behind those records?
- How many people meet a chosen definition of active participation today?
Each question produces a different result. For planning recruitment, the active-person view is usually the useful one. For reconciling an official register, the authorisation count may be the correct measure. For analysing the historic reach of the market, the broader person universe can matter.
The danger begins when a system answers one of those questions while presenting the result as if it answered all three.
Why AI increases the need for definitions
Traditional analysis was often slow enough that an analyst had time to notice an odd result. AI can generate a fluent answer immediately. Speed is valuable, but it also makes a plausible mistake easier to trust.
Suppose one advisor receives six product authorisations at a new company. A naive query can report six arrivals. The calculation is internally consistent, but the management conclusion is wrong. The same problem appears in departures, tenure, company transitions and product specialisation.
This is why Emable separates the layers:
- The original regulatory record is preserved.
- Records are resolved to a person identity where the evidence supports it.
- Product-company authorisation changes are kept distinct from inferred physical career moves.
- Every metric carries a declared population, observation date and definition.
- Low-confidence matches remain visible instead of being silently converted into facts.
The AI works on top of those rules. It does not get to invent them.
The historical layer matters too
Regulatory data has a biography. Systems change, historical records are imported, names are corrected and reporting practices evolve. A sudden spike in a time series may describe a real market event, or it may describe the day a historical stock was loaded into a new system.
An experienced analyst learns these discontinuities. A reusable intelligence product must encode them so the knowledge does not live only in one person’s memory.
For Emable, that means maintaining explicit quality checks for:
- unexpected changes in record volume;
- duplicate or conflicting identities;
- authorisations with impossible or incomplete date relationships;
- changes in product and role codes;
- gaps between record-level and person-level movement.
These checks are not glamorous. They are also more important than another layer of generated commentary.
What decision-grade regulatory intelligence looks like
A useful answer should make its scope inspectable. A manager should be able to see whether a figure represents people, registrations or product-company relationships. They should know the date, geography and activity criteria used. If an estimate depends on identity resolution, the method and uncertainty should be stated.
This does not make the product less sophisticated. It makes the sophistication accountable.
The most trustworthy AI systems in finance will not be the ones that always sound certain. They will be the ones that know which question they answered, retain the evidence behind it and show where the boundary of confidence lies.
That is the difference between generating a number and building intelligence.
Methodology and limitations
This article describes Emable's internal person-resolution and authorisation-history methodology. Counts are estimates derived from regulatory records and depend on the definition and observation date used.