Why pharma's next regulatory advantage is a master data problem, not an analytics one.
In March 2026, the FDA quietly adopted ICH M14, the international guidance on the design and reporting of non-interventional studies that use real-world data for safety assessment. The adoption did not generate the headlines that an AI announcement would have. It should have. With that step, the agency completed the transition that began with the 21st Century Cures Act in 2016 and accelerated through the Advancing Real-World Evidence Program in 2022: real-world evidence is no longer an experimental supplement to randomised trials. It is a first-class pathway for label expansions, post-marketing commitments, and safety surveillance, with its own technical standards, its own data-quality bar, and a peer-reviewed body of regulatory precedent the agency now references in its own decision letters.
For pharma, this is the most consequential regulatory shift since the introduction of accelerated approval. It changes what evidence looks like, who generates it, and how long a label expansion takes from first conversation to approval. But the shift has also exposed a structural problem the industry has been deferring for a decade: the data foundation underneath RWE is, in most companies, not ready for the standard the FDA is now applying. The companies that recognise this and treat it as a master data problem rather than an analytics problem will get their labels expanded faster. The companies that treat it as an analytics problem will be doing the same work twice, slowly.
This piece is about why.
The state of RWE in 2026
The volume of RWE-supported submissions has grown steadily but unevenly. A peer-reviewed analysis of 218 FDA labeling expansions granted between January 2022 and May 2024 found that roughly a quarter of them — between 23 and 28 percent depending on the year — were supported in part by real-world evidence. The proportion is rising, but the distribution by therapeutic area is far from even.

Oncology dominates the picture, accounting for nearly forty-four percent of all RWE-supported label expansions. The reason is straightforward: oncology has the deepest real-world data infrastructure in the industry, with platforms like Flatiron Health, Tempus, Syapse, and Komodo Health building structured longitudinal datasets across hundreds of community and academic sites. Infectious disease and dermatology follow at roughly nine and seven percent, with cardiovascular, neurology, endocrinology, and rheumatology each in the four-to-six percent range. The "other / mixed" category at twenty-one percent reflects the long tail of rare-disease and special-population approvals where RWE often plays a decisive role precisely because randomised trial enrolment is impractical.
The market following this regulatory shift is large and growing. Estimates vary depending on how broadly the analysts define the category, but every reading shows the same shape. The narrow real-world data market is projected at roughly $2.7 billion in 2026, growing at a 14-15% compound annual rate. The broader real-world evidence solutions market, which includes the analytics, services, and platform layers built on top of the data, is sized at over $20 billion in 2026 and projected to nearly triple by the early 2030s. Tempus alone reported Q1 2026 revenue of $348 million, up 36 percent year-on-year, with its data-and-applications segment growing at over 40 percent. The platforms are growing because the spend is growing. The spend is growing because the regulatory environment now justifies it.
Where the evidence actually comes from
The phrase "real-world data" makes RWE sound simpler than it is. In practice, the evidence the FDA accepts is assembled from a federation of data sources, none of which is sufficient on its own.

Electronic health records contribute the largest share of the underlying signal — roughly a quarter of what a regulator-grade RWE study draws from — because EHRs hold the clinical detail that claims data cannot. Administrative claims contribute almost as much, because claims hold the longitudinal coverage, the prescription-fill history, and the economic outcomes that EHRs cannot reliably reconstruct. Disease and patient registries add another fourteen percent, providing the disease-specific clinical depth that neither EHR nor claims can match for the conditions in scope. Connected devices and wearables, patient-reported outcomes, lab and diagnostic data, mortality and vital records, and increasingly genomic and molecular data round out the fabric.
The composition matters because it determines what the study can and cannot say. A claims-only study can describe utilisation and economic outcomes with confidence but says little about clinical response. An EHR-only study describes clinical response but is silent on what happened to the patient between encounters, when they switched providers, or whether they filled the prescription. A registry-only study has depth but suffers from selection bias the FDA increasingly scrutinises. The strongest regulatory submissions are the ones that combine multiple sources into a single longitudinal patient record — and the weakest ones, the ones that get pulled back for protocol revisions, are almost always the ones built from a single source.
Why one source is never enough
The diagnostic value of each RWD source varies by what the study is trying to evidence. The matrix below, distilled from the dimensions that actually appear in FDA review letters, shows where each source is strong, partial, weak, or absent.

The matrix makes a simple point in a way that is harder to argue with than prose. EHR is authoritative for diagnosis detail but partial on long-term follow-up. Claims is authoritative for treatment-and-adherence patterns and for economic outcomes but weak on clinical detail. Registries are strong on clinical and follow-up dimensions for the conditions they cover, but their patient-identity link to the broader RWD ecosystem is partial at best. Connected devices are authoritative for real-time signal and partial on almost everything else. Mortality records are authoritative for one thing — survival — and absent on the rest.
The implication for any pharma company planning a label-expansion strategy that depends on RWE is direct. A defensible RWE submission requires data from at least three of these sources, linked at the patient level, with each source contributing to the dimensions where it is strongest. The companies that have built the linkage infrastructure can move from study design to submission in months. The companies that have not are still hunting for the third source twelve months into the project.
The patient identity problem nobody named in the press release
The hardest part of building this multi-source RWE fabric is not the analytics. It is the patient identity layer underneath it.

Each RWD source captures a partial picture of patient identity. EHR systems hold the strongest demographic and clinical detail — full name at 92 percent coverage, date of birth at 94 percent, diagnoses at 96 percent — but their master patient ID is internally consistent only within the deploying health system, and the cross-EHR linkage rate against the broader ecosystem is meaningfully lower, around 70 percent in our composite view. Claims data has stronger demographic alignment for billing reasons but limited contact information and almost no lab or genomic detail. Disease registries have deep clinical coverage for their patient population but lower demographic completeness because patients are often enrolled with minimal personally identifying information by design. Mortality and vital records have authoritative survival data but essentially no clinical signal.
Linking these sources at the patient level — without breaching privacy, without losing fidelity, and with a chain of consent and authorisation the FDA can audit — is the work that makes RWE submissions defensible. It is also the work that almost no pharma company has built once and built well. The industry has historically treated patient identity resolution as a one-off project for each study. The companies that are pulling ahead in 2026 are treating it as durable infrastructure that every study reuses, and the regulatory acceptance follows.
The cost economics tell the same story from a different angle. Industry analyses now place the cost of healthcare data linkage services frequently above the cost of the underlying data licences themselves. The data is increasingly commoditised. The linkage is not. The companies that own a reusable, governed, audit-ready linkage capability — typically anchored on a patient master data platform integrated with tokenisation services like Datavant, HealthVerity, or the equivalent — are the ones that submit studies on schedule. The companies that contract for linkage on a per-study basis are the ones that explain to their executive committee why the H1 submission slipped to H2.
What this means for the pharma data foundation
The architectural answer is recognisable. RWE at scale requires the same data-foundation pattern that every other modern enterprise AI use case requires, applied to the specific entities and sources pharma operates over.
A trusted patient master, governed at the brand level, with cross-source identity resolution and a consent and authorisation record that survives external audit. A drug master that ties the company's commercial portfolio to industry-standard reference data — RxNorm, NDC, IDMP — with the level of attribute completeness FDA submissions require. A provider master that resolves the same physician across EHR, claims, and registry sources. A study master that tracks every protocol, every data source, every linkage operation, and every analytic output with versioned lineage from the source system to the regulator-facing exhibit.
This is master data management applied to life sciences, with the same discipline Apptad has been deploying across Reltio, Informatica, and STIBO for other domains, configured for pharma's specific entity model. The work is unglamorous. It is also the precondition for everything the agency now expects RWE submissions to demonstrate.
The companies that have invested here in 2024-2025 are the ones whose RWE pipelines are running at six-to-nine-month cycle times. The companies that have not are still running at eighteen-to-twenty-four months, with each study reinventing the linkage and identity work the previous study should have produced as reusable infrastructure.
A 90-day diagnostic for pharma data and regulatory leaders
The first thirty days are a study-portfolio inventory. For every RWE-dependent submission in the pipeline — current, planned, and aspirational — document the patient population, the data sources required, the linkage operations the study depends on, and the lineage and consent chain the regulator will be asked to inspect. The output is a heatmap of where the same data foundation could serve multiple studies and where each study is asking the organisation to do the work for the first time.
The second thirty days are a foundation assessment. Inventory the patient master, drug master, provider master, and study master as they exist today. Assess each against the four standards the FDA now applies in RWE review: completeness across the relevant population, accuracy of the cross-source linkage, freshness against the study reference window, and auditability of the lineage and consent record. The output is a short list of foundation gaps that, if closed, would unblock the largest share of the study portfolio.
The third thirty days are an operating-model decision. The choice is whether to continue treating the foundation work as a per-study cost line — buying linkage and identity resolution from a vendor for each study, with no reuse — or to build it as durable enterprise capability, owned by a pharma master data function, governed centrally, and serving every study and downstream commercial use case. The economic case for the second path has been clear in our practice for two years. The regulatory case, with ICH M14 now adopted, is now decisive.
The reframing
RWE in 2026 is no longer a question of whether the FDA will accept real-world data. The agency has answered that question. The question is whether your data foundation is ready to produce evidence at the standard the agency is now applying. Most pharma data foundations were designed for trial data, where the source of truth is the eCRF and the patient is the trial participant. They were not designed for the messy, federated, multi-source, longitudinal world that real-world evidence requires. The companies that recognise the gap and close it deliberately will compete on a faster regulatory clock than the companies that continue to treat each RWE submission as a bespoke data engineering project.
Apptad partners with pharma CIOs, CDOs, heads of regulatory affairs, and clinical operations leaders to design and stand up the data foundation that makes RWE at scale possible — patient, drug, provider, and study master data; integrated linkage and consent infrastructure; and the operating model that turns the foundation into a reusable competitive asset across submissions. If your H1 2026 has made it clear that the agency is moving faster than your foundation can support, H2 is the right window to close the gap. That is the conversation worth having.



