Real-world data licensing is expensive: hundreds of thousands to millions of dollars a year. A data provider's presentation can make a dataset look ideal: hundreds of millions of patients, years of follow-up, and a long list of clinical variables. Your study, however, depends on a much smaller set of requirements. Can you identify the right patients, observe the treatment decision, measure the outcome, and account for important differences between groups?
Consider a hypothetical oncology study comparing two therapies. A database may contain both drugs and thousands of patients with the diagnosis. Yet the analysis can still be derailed by incomplete prior-treatment history or an outcome recorded inconsistently across sites.
Before committing to a license, translate the research question into evidence requests the provider can answer. These five questions offer a starting point.
1. How many patients actually meet our study requirements?
Start with the population you intend to study, including disease subtype, treatment setting, geography, and calendar period. Then ask how many patients remain after applying the criteria your analysis needs.
For example, a count of patients with ovarian cancer does not establish how many have a particular histology, received a specified treatment sequence, and have sufficient baseline information. Each additional requirement changes both the number of eligible patients and who they represent.
Ask the provider for: a staged cohort count showing the effect of each proposed inclusion criterion, with definitions and dates. Where a criterion cannot be evaluated, ask that it be labeled unknown rather than silently omitted.
Use this exercise to test feasibility and identify selection concerns. A large starting population is no guarantee that the final cohort supports the planned analysis.
2. Can we measure the endpoint our question requires?
A variable name is only the beginning. Establish what an endpoint means in the dataset, where it comes from, and how it was derived or validated. FDA's July 2024 guidance on EHR and claims data addresses the definition, ascertainment, and validation of study outcomes.1
In our hypothetical oncology study, a field labeled “progression” could reflect a clinician's assessment, a curated interpretation of notes, or an algorithmic proxy. Examine each against the intended outcome. Do not assume the label establishes equivalence to a trial endpoint.
Ask the provider for: the operational definition, source information, derivation method, available validation evidence, and missingness within the proposed cohort. If a proxy is necessary, decide whether it still answers the original question and document the change.
3. Can we reconstruct the timeline the design needs?
Draw a simple patient timeline before evaluating the data: eligibility assessment, baseline measurement, treatment initiation, and follow-up. Identify which dates must be available and what each recorded date represents.
For a treatment comparison designed to emulate a trial, eligibility, treatment assignment, and the start of follow-up must align. Misalignment can introduce selection or immortal-time bias.2
One practical request: ask the provider to walk through synthetic patient timelines showing how it identifies treatment initiation and distinguishes prior treatment from a new episode. The purpose is to make the proposed rules inspectable before they become embedded in an analysis.
Ask the provider for: definitions of treatment and observation dates, the availability of baseline history, and how gaps or the end of observable follow-up are identified.
4. Are important confounders measured when we need them?
For a comparative study, identify factors that could influence both treatment choice and the outcome. Determine whether the dataset measures them with suitable definitions and timing. FDA's guidance discusses covariate ascertainment and validation, including confounders and effect modifiers.1
For example, a study team might need pretreatment disease severity or functional status. A value recorded only after treatment begins may not serve the intended baseline role.
Ask the provider for: a feasibility table showing the definition, measurement window, and completeness of each critical covariate in each treatment group. Assess whether gaps concentrate in particular sites, periods, or patient groups.
Decide with the study methodologist whether the remaining limitations require a different design, additional data, or a narrower question. A promise of later statistical adjustment does not resolve a feasibility problem.
5. What can we verify before making the commitment?
Some assessments must proceed without sponsor access to patient-level records. A 2026 study of registry-based post-authorization safety studies examined that situation and found gaps in the evaluated assessment tools. The authors proposed quantitative data-quality indicators to complement documentation-based assessment.3
Agree on a small, study-specific feasibility package. Where access is restricted, the provider may be able to run agreed checks and return aggregate results, subject to its governance rules.
Ask the provider for: evidence supporting the critical assumptions, a description of unresolved gaps, and an explanation of what further assessment would require. Agree on how an unmet requirement would affect the next step before committing to the full study.
Take this checklist to your next provider discussion

The outcome of these discussions should be a documented decision: proceed, proceed with specified limitations, revise the design, or evaluate another source. That decision is most useful when it connects each important data limitation to its consequence for the study.
Evaluating a dataset for an RWE or HEOR project? Talk with Polygon Health Analytics about your research question and data requirements.
References
1. U.S. Food and Drug Administration. Real-World Data: Assessing Electronic Health Records and Medical Claims Data To Support Regulatory Decision-Making for Drug and Biological Products. Guidance for Industry. July 2024.
2. Hernán MA, Sauer BC, Hernández-Díaz S, Platt R, Shrier I. Specifying a target trial prevents immortal time bias and other self-inflicted injuries in observational analyses. J Clin Epidemiol. 2016;79:70-75.
3. Dobay P, Sabidó M. Evaluating Data Quality by Proxy: Can We Evaluate All Dimensions of the European Medicines Agency Data Quality Framework for Registry-Based Post-Authorization Safety Studies? Pharmacoepidemiology and Drug Safety. 2026;35.