Almost everyone in clinical research says “data-driven.” Then the data underneath turns out to be inferred — therapeutic-area averages, registry estimates, past trial counts, a reputation list of the usual names. We run site selection on something else entirely: real, patient-level evidence. So the sites you pick are the sites that actually enroll.
Real, patient-level data is the de-identified record of what actually happened to real patients — across every institution that treated them, over time. Not a survey. Not a therapeutic-area estimate. Not who published the most papers. It is the counted reality of which protocol-qualifying patients exist and which investigators reach them, down to a number.
The industry standard is to infer. Start with a therapeutic-area average, layer on a registry estimate and a list of investigators known by reputation, and present it with confidence. It looks like evidence. It enrolls like a guess — which is why so many sites miss their targets and trials scramble to add rescue sites to catch up.
Both answer “how many patients can this site enroll?” Only one of them is counted.
Illustrative figures. The point isn’t the number, it’s that one is a guess and one is counted, per site, before you commit a dollar.
Everyone brings a site list. We bring the evidence underneath it — the part a sponsor can’t argue with.
For every candidate site, the actual number of patients who meet your protocol’s criteria — under an investigator’s direct care and across their institution. Counted, not inferred.
A shortlist of consistently high-performing principal investigators, scored on a measured enrollment record — not a reputation list of the usual academic names.
The whole patient journey across the entire U.S. — every institution, never a single network’s snapshot. The most complete real-world evidence base in the country.
The same protocol, run on real evidence instead of inference, produces a different, better site list — and a trial that behaves:
We hand you the finished evidence — the ranked, defensible site list and the reasoning behind it. We do not pass raw patient data through to you, and it never trains anyone else’s model. It is de-identified by design and confidential to your program. That boundary isn’t a limitation; it’s part of why sponsors trust the result.
In a recent rare-disease program, modeling the protocol on real, patient-level data nearly tripled the eligible pool — then became a ranked list of the sites that could actually reach those patients, on one evidence base.
Six questions to ask anyone who tells you their site selection is “data-driven.”
Is there a real, protocol-qualifying patient count per site — down to a number — or a therapeutic-area estimate?
Are investigators ranked on a measured enrollment track record, or on reputation?
Is the data national and longitudinal, or a single network’s snapshot?
Does it follow the patient across institutions, or stop at one system’s walls?
Can they explain why each site made the list — and why some marquee names didn’t?
Is it de-identified and confidential to your program? Real evidence and real privacy are not a trade-off.
What is real patient data in clinical trials?
Real patient data is the de-identified record of what actually happened to real patients, across every institution that treated them, over time. In site selection it means the counted number of protocol-qualifying patients an investigator can reach — not a therapeutic-area estimate or a registry sample.
How is patient-level data different from registry or inferred estimates?
Inferred estimates start from a sample or an average and project outward. Real, patient-level data counts the actual qualifying patients per site and follows each patient across institutions rather than stopping at one system’s walls. One is a likelihood; the other is a number.
Do you share the patient data with us?
No. We deliver the outcome — a ranked, defensible site list and the reasoning behind it — not the raw patient data. It is de-identified by design, confidential to your program, and never trains anyone else’s model.
Why is real data better than inferred data for site selection?
Because a site list built on real, counted patients enrolls, and one built on inference often doesn’t. Real data means first-patient-in faster, a ramp that holds, and fewer amendments and rescue sites from discovering an enrollment gap months too late.
Bring us your protocol. We’ll come back with the sites and investigators that will actually enroll it — each one backed by real, patient-level evidence.