R RegData
All articles
IDMP eCTD 4.0 SPOR Controlled Vocabulary Veeva Vault RIM MHRA EMA Regulatory Data Architecture Regulatory Information Management (RIM)

Why Regulatory Data Still Doesn't Talk to Itself

Over the past month, a recurring theme has cut across otherwise separate pieces of work our consultants have undertaken: Cross-regulatory agency submission classification design, an IDMP interoperability proposal, and a review of how a publishing platform handles controlled vocabulary. Different projects, same underlying fault line — regulatory data architecture is still fighting a fragmentation problem that no single vendor or standard has actually solved.

The Controlled Vocabulary Problem: Why Regulatory Data Still Doesn't Talk to Itself

Over the past month, a recurring theme has cut across otherwise separate pieces of work our consultants have undertaken: Cross-regulatory agency submission classification design, an IDMP interoperability proposal, and a review of how a publishing platform handles controlled vocabulary. Different projects, same underlying fault line — regulatory data architecture is still fighting a fragmentation problem that no single vendor or standard has actually solved.

One term, two meanings

The clearest example came from building a master picklist architecture for a regulatory system that we are implementing for one of our clients. "Procedure Type" looked like a single field until it was pressure-tested against both regional /country level values for “Regulatory Objectives” (the application pathway, set once) and the global “Regulatory Events” (the discrete submission activity, set per event). Trying to serve both from one flat list conflates two genuinely different axes — pathway versus activity — and breaks the moment an agency's structure doesn't map cleanly onto EMA's Centralised/Decentralised/MRP model.

That mapping problem gets worse cross-agency. FDA and Health Canada aren't procedure-choice authorities at all — NDA, ANDA, BLA, NDS, ANDS are legal-basis concepts wearing a procedure-type costume. TGA and Swissmedic add their own categories again. Expedited status (Priority Review, Breakthrough Therapy, FTP) doesn't belong in any of these lists as a mutually exclusive value — it's an orthogonal decision bolted onto whichever regulatory pathway applies. A vocabulary designed around one agency's mental model quietly imports that agency's assumptions everywhere else it's reused.

Controlled vocabularies are versioned locally, not exchanged globally

A separate review of one of the eCTD Publishing systems surfaced a gap that's easy to miss until you go looking for it: there's no publicly documented, named mechanism for controlled vocabulary exchange between a publishing engine and a RIM system holding structured regulatory / submission data. Both systems maintain their own CV module and their own versioning cadence, and both nominally draw from the same upstream sources (EMA SPOR, ICH/regional eCTD codelists) — but "both subscribe to the same standard" is not the same as "both stay synchronised." In practice, CV alignment between authoring and publishing layers tends to be custom-built or manually reconciled per client rather than a standard, contractually guaranteed connector. For any organisation running a RIM-to-publishing pipeline, that's a data-quality risk hiding in plain sight, not a solved integration.

The common thread

None of these are separate problems. They're the same problem at different resolutions:
- Field Level: a single controlled term asked to represent two different concepts for two different consumers.
- System Level: two systems each maintaining their own version of "the same" vocabulary with no guaranteed sync.
- Standard level: a global data model too complex for most regulators to adopt in full, leaving the exchange layer to be defined by whoever moves first.
- Policy Level: legal frameworks going live before the vocabulary and guidance that should underpin them.

Practical Implications

The practical implication for anyone designing regulatory data architecture is the same at every level: treat controlled vocabulary design as a first-class architectural decision, not a configuration afterthought. Split axes that look similar but aren't (pathway vs. activity, legal basis vs. procedure). Don't assume two systems referencing the same external standard are actually synchronised — verify or contract for it explicitly. And where a global standard is too heavy to adopt wholesale, an intentionally lightweight, aligned-not-complete subset — scoped to the highest-friction exchanges rather than full coverage — is often the more realistic path to interoperability than waiting for full IDMP adoption to arrive.

Our Tookits

Our Controlled Vocabulary (CV) Management toolkit enables our clients to manage these complexities and build a strong Data governance process to set them on path to adopting eCTD 4.0 and IDMP alignment

Why not talk to our experts, who have years of experience supporting RIM Implementation

Back to all articles