Clinical Trial Protocol Complexity 2026: Endpoints, Multicenter Countries and Study-Build Workload on ClinicalTrials.gov
TL;DR
Takeaway: ClinicalTrials.gov and AACT store multipliers — specified outcomes, design groups, countries, facilities — not hours to build an EDC. Across 595,630 studies the median NCT has 4 specified outcomes and 1 country; 42,756 studies (7.2%) list two or more countries. The Phase 3 industry interventional cohort (n = 24,676) is a different object: p90 = 20 outcomes and 17 countries.
A study-build lead who treats every NCT as the same database is staffing a single-site observational study and a 17-country Phase 3 from the same template. The public file will not tell you how many visits, CRF pages, edit checks or translations the build contains. It will tell you which band you are in before you commit writers and UAT budget.
Three computed facts set the scale.
On the 25 July 2026 ClinicalTrials.gov study file joined to AACT calculated values (1 August 2026): median 4 specified outcomes, p90 = 13, p99 = 36, max 2,334, mean 6.23. Design groups: median 2, p99 = 7, max 63. Half of registered studies have four or fewer specified outcomes. The mean is a tail statistic. Quote the percentiles [1][2][3].
Countries are not sites. Non-removed country rows: median 1, p90 = 1, p95 = 2, p99 = 12, max 59. 42,756 of 595,630 studies (7.2%) have two or more countries. Facilities: median 1, p90 = 8, p95 = 22, p99 = 102, max 3,510. A country count used as a site count is the wrong grain [1][2].
Phase 3 industry interventional studies (phase string contains PHASE3, sponsor class INDUSTRY, study type INTERVENTIONAL; n = 24,676) sit on a different surface: outcomes median 6, p90 = 20, p99 = 52, max 214, mean 9.4; countries median still 1, but p75 = 7, p90 = 17, mean 5.45, max 59. Observational studies (n = 139,061) have median 3 outcomes. The 2,334-outcome maximum in the all-study file is an observational row. It must not set the quote for a Phase 3 industry build [2][3].
This paper is a sizing workpaper. It is not a “protocols are getting more complex” essay. Adjacent EClinCloud Deep Research covers cross-registry identity (24 August 2026) and COA language in public outcome rows (28 August 2026). Those papers answer different questions. This one asks which observable public-protocol objects multiply the study-build surface, and how to quote a build without inventing an hour claim.
One NCT is not one workload
Takeaway: The 25 July 2026 ClinicalTrials.gov file holds 595,630 NCT rows. INTERVENTIONAL 454,529 versus OBSERVATIONAL 139,061, and INDUSTRY 131,151 versus OTHER 425,599, are already two different build surfaces — before anyone counts endpoints or countries.
An NCT number is a study-record identifier. ClinicalTrials.gov’s own glossary defines a study record as a protocol summary — recruitment status, eligibility, contacts, and in some cases summary results — not a configured database [4][5]. Protocol registration data elements, submitted through the Protocol Registration and Results System, capture what the sponsor entered about design, not what a data manager later built [1][6]. Treating “one NCT = one EDC study of typical size” is a unit error at the first row.
The 25 July 2026 downloadable study file is the denominator for every per-study figure in this paper: 595,630 unique NCT rows [3]. Study type:
| Study type | Studies | Share of 595,630 |
|---|---|---|
| INTERVENTIONAL | 454,529 | 76.3% |
| OBSERVATIONAL | 139,061 | 23.3% |
| EXPANDED_ACCESS | 1,059 | 0.2% |
| BLANK | 981 | 0.2% |
ClinicalTrials.gov study records, 25 July 2026 — EClinCloud analysis, accessed August 2026 [3].
Interventional records are three-quarters of the file. Observational records are still 139,061 studies — larger than the entire Phase 3 industry interventional cohort computed later. An observational registry with three specified outcomes and one country is a real NCT. A randomized, parallel, masked Phase 3 with twenty specified outcomes and seventeen countries is also a real NCT. The identifier does not encode the difference.
Sponsor class is the second cut. ClinicalTrials.gov classifies the lead sponsor; INDUSTRY is not a synonym for “commercial quality,” and OTHER is not a synonym for “academic small”:
| Sponsor class | Studies | Share of 595,630 |
|---|---|---|
| OTHER | 425,599 | 71.5% |
| INDUSTRY | 131,151 | 22.0% |
| OTHER_GOV | 15,780 | 2.6% |
| NIH | 11,559 | 1.9% |
| NETWORK | 4,990 | 0.8% |
| FED | 4,914 | 0.8% |
| BLANK | 981 | 0.2% |
| INDIV | 559 | 0.1% |
| UNKNOWN | 94 | <0.1% |
| AMBIG | 3 | <0.1% |
ClinicalTrials.gov study records, 25 July 2026 — EClinCloud analysis, accessed August 2026 [3].
Industry is 22.0% of registered studies. Most NCT rows are OTHER. A complexity score computed on the whole file is a score of the OTHER-heavy mass, not a score of the industry Phase 3 a vendor is being asked to build. The useful comparison is not “industry versus the world” as a quality ranking. It is “which slice of the file matches the protocol on my desk.”
Phase makes the same point from another column. Of 595,630 studies, phase NA is 232,760 and phase BLANK is 141,220. A clean PHASE3 label is 41,964; PHASE2 is 64,983; PHASE1 is 48,269; PHASE4 is 35,564; PHASE2;PHASE3 is 7,547 [3]. Most registered studies do not have a Phase 3 label. Saying “Phase 3 protocols are complex” without naming the sponsor class and study type is a slogan about a minority of the file.
Status is not workload either. COMPLETED 325,239; UNKNOWN 94,294; RECRUITING 65,408; TERMINATED 34,011 [3]. A completed observational study and a recruiting multinational Phase 3 can share a status word. Status answers “where is this record in the registry lifecycle.” It does not answer “how many forms will we build.”
Posted results are a third object. 79,339 studies in this snapshot have results posted [3]. Results data elements — participant flow, baseline characteristics, outcome measures, adverse events — are defined separately from protocol registration [7]. Protocol-specified outcomes in AACT design_outcomes are the planning surface. Posted result tables are what was submitted after the fact. Do not size a study build from a results-posting rate.
The decision that changes after this section is simple. Before quoting build effort, read study type, sponsor class and phase from the NCT. If the protocol is industry, interventional and Phase 3, you are not in the median NCT. You are in a 24,676-row cohort with its own percentiles, computed below. If it is observational, you are in a 139,061-row cohort whose median outcome count is three. One identifier, two (actually many) workloads.
ICH E6(R3) asks that computerized systems fit the trial [8][9]. The public row can tell you the trial’s registered type. It cannot tell you whether the system that will capture it is proportionate. That gap is the rest of this paper.
Outcomes as the first multiplier
Takeaway: Specified-outcome count is the first public multiplier. Median 4, p90 13, p99 36, max 2,334, mean 6.23 on 595,630 studies. AACT holds 3,716,648 design_outcome rows — 1.17 million primary, 2.34 million secondary. That mix is a naming split, not a COA taxonomy.
ClinicalTrials.gov defines an outcome measure as a planned measurement in the protocol: for trials, an effect of the intervention; for observational studies, a measurement used to describe patterns or associations. Types include primary and secondary [1][4]. The definition is a protocol object. It is not a CRF page, not an edit check, and not an instrument version.
AACT stores those protocol objects as rows. The 1 August 2026 AACT pipe files contain 3,716,648 design_outcomes rows and 596,685 calculated_values rows. Joined to the 595,630-NCT ClinicalTrials.gov file, specified outcomes per study are:
| Statistic | Specified outcomes per study |
|---|---|
| n | 595,630 |
| Minimum | 0 |
| Median (p50) | 4 |
| p75 | 8 |
| p90 | 13 |
| p95 | 18 |
| p99 | 36 |
| Maximum | 2,334 |
| Mean | 6.23 |
ClinicalTrials.gov / AACT — EClinCloud analysis, accessed August 2026 [2][3]. AACT design_outcomes row counts per study match the calculated-values total in this join: same median, same p90, same max.
Specified outcomes per study: median, p90 and p99
The median NCT specifies 4 outcomes; the Phase 3 industry interventional median is 6 and p90 is 20. Observational median is 3. Maxima (2,334 observational; 214 in the Phase 3 industry cohort) are omitted so the percentiles remain readable — they must not set a mean-based quote.
Read the table as a staffing filter, not as a complexity trend.
Half of registered studies specify four or fewer outcomes. Three-quarters specify eight or fewer. The study at p90 specifies 13. The study at p99 specifies 36. The study at the maximum specifies 2,334. That maximum is a real row in the file. It is also an outlier that must not be allowed to set the mean narrative: mean 6.23 versus median 4 is the tail talking. A quote that uses the mean as “typical” has already overstated the typical NCT and understated the tail.
Measure text is blank on 9 of 3,716,648 rows. The outcome field is filled. What it is filled with is names. Primary versus secondary in this file is a registry labelling split: 1,174,906 primary rows, 2,335,625 secondary, 206,117 other. Secondary rows are 62.8% of the outcome table; primary rows are 31.6% [2]. A secondary line can be a serious endpoint or a sparse exploratory name. The type field does not tell you which. It tells you how the sponsor labelled the row.
That labelling mix is as far as this paper goes into outcome kind. An adjacent Deep Research piece (28 August 2026) asks whether FDA PFDD Guidance 3 taxonomy — PRO, ClinRO, ObsRO, PerfO — appears in these same rows. It almost never does. This paper does not retell that argument. For study-build sizing, the operational fact is the count of specified outcomes and the primary/secondary split as a naming mix, not a validation file [10].
What the count is good for: a first band. A protocol with 3 specified outcomes and a protocol with 20 specified outcomes are different CRF-and-edit-check surfaces even if every other public field is equal — and they are not equal, as the later sections show. What the count is not good for: converting 13 outcomes into 13 forms, or 13 forms into N hours. One specified outcome can be a single binary field. Another can be a multi-instrument composite collected at every visit in three languages. The public row does not hold visit structure.
Action after this section: pull the outcome count from AACT or from the protocol before quoting build effort. If the NCT is at or above p90 (13 outcomes in the all-study file; 20 in the Phase 3 industry interventional file), do not quote it as a median study. If it is the 2,334-outcome row, do not let it anywhere near a mean-based rate card.
Groups and masking
Takeaway: Design groups are a second multiplier — median 2, p99 7, max 63 — not a small-study guarantee. EXPERIMENTAL is 532,765 of 1,097,076 group rows. RANDOMIZED is 299,465 of 591,898 design rows; PARALLEL 271,104. Masking NONE is 249,894 design rows. NONE is not simpler ethics.
AACT design_groups holds 1,097,076 rows. Per study, joined to 595,630 NCTs: median 2 groups, p75 = 2, p90 = 3, p95 = 4, p99 = 7, max 63, mean 1.84 [2]. Most registered studies are two-group objects on this field. The tail is real. A 63-group row is not a typical Phase 3, and it is not fictional.
Group type is a row vocabulary, not a study census. Of 1,097,076 group rows:
| Group type | Group rows |
|---|---|
| EXPERIMENTAL | 532,765 |
| ACTIVE_COMPARATOR | 201,248 |
| BLANK | 169,495 |
| PLACEBO_COMPARATOR | 80,404 |
| NO_INTERVENTION | 57,710 |
| OTHER | 43,035 |
| SHAM_COMPARATOR | 12,419 |
AACT design_groups, 1 August 2026 — EClinCloud analysis, accessed August 2026 [2].
AACT design-group rows by type (not a study census)
EXPERIMENTAL is 532,765 of 1,097,076 group rows (48.6%). These are group rows, not studies: a two-arm trial contributes two rows. BLANK (169,495) is a missing type label, not a missing arm.
EXPERIMENTAL is 48.6% of group rows. ACTIVE_COMPARATOR is 18.3%. BLANK is 15.5% — a missing type label, not a missing arm. PLACEBO_COMPARATOR is 80,404 rows. These counts are rows. A two-arm study contributes two rows. A factorial study contributes more. Do not read 532,765 as 532,765 studies with an experimental arm.
Allocation and intervention model live on AACT designs (591,898 rows — one design record per study in this extract, short of the 595,630-NCT file by the snapshot gap):
| Allocation | Design rows |
|---|---|
| RANDOMIZED | 299,465 |
| BLANK | 142,236 |
| NA | 102,536 |
| NON_RANDOMIZED | 47,661 |
| Intervention model | Design rows |
|---|---|
| PARALLEL | 271,104 |
| BLANK | 143,136 |
| SINGLE_GROUP | 122,818 |
| CROSSOVER | 35,841 |
| SEQUENTIAL | 13,049 |
| FACTORIAL | 5,950 |
AACT designs, 1 August 2026 — EClinCloud analysis, accessed August 2026 [2].
Allocation on AACT design rows
RANDOMIZED is 299,465 of 591,898 design rows (50.6%). NA (102,536) is often the honest value for single-group or observational designs. BLANK (142,236) is a fill limitation, not proof the study had no design. Randomization supply remains an RTSM object, not an EDC hour.
RANDOMIZED is 50.6% of design rows. PARALLEL is 45.8%. SINGLE_GROUP is 122,818. NA allocation is 102,536 — often the honest value for a single-group or observational design, not a quality defect. BLANK on 142,236 allocation rows and 143,136 model rows is a fill limitation of the submitted record. It is not proof the study had no design.
Masking is the field people misuse as a proxy for “how hard will ethics and blinding logistics be”:
| Masking | Design rows |
|---|---|
| NONE | 249,894 |
| BLANK | 141,939 |
| SINGLE | 68,721 |
| DOUBLE | 61,123 |
| QUADRUPLE | 40,415 |
| TRIPLE | 29,806 |
AACT designs, 1 August 2026 — EClinCloud analysis, accessed August 2026 [2].
NONE is the largest labelled bucket: 249,894 rows, 42.2% of design records. Open-label interventional trials are common. Open-label is not “simpler ethics.” An unblinded oncology study with independent review, extensive safety reporting and a single experimental arm can be a heavier operational object than a double-blind two-arm allergy study. The masking field records how the sponsor described blinding. It does not record IRB workload, DSMB structure, or whether endpoint review is centralized.
Primary purpose, for context rather than as a complexity score: TREATMENT 289,414; BLANK 143,365; PREVENTION 48,165; SUPPORTIVE_CARE 25,802; BASIC_SCIENCE 21,640; DIAGNOSTIC 19,355 [2]. TREATMENT is 48.9% of design rows. The rest of the file is not a residual of “other Phase 3.”
Randomization and trial supply are not this paper’s object. A RANDOMIZED + PARALLEL + PLACEBO_COMPARATOR study still needs an RTSM design that the registry will not hold — kit types, stratification, resupply. That work is RTSM, not EDC [11]. Site activation, milestone tracking and monitoring plans are CTMS, not EDC [12]. The group and masking fields help you see that the protocol is not a single-arm open registry. They do not collapse three products into one build estimate.
Action after this section: count groups from the protocol or from AACT before assuming a two-arm template. If masking is NONE, do not discount the ethics and endpoint-review surface. If allocation is RANDOMIZED, budget RTSM as a separate object.
Countries are not sites
Takeaway: Country rows are not investigational sites. Non-removed countries per study: median 1, p90 1, p95 2, p99 12, max 59. Only 42,756 of 595,630 studies (7.2%) have two or more countries. Facilities: median 1, p90 8, p95 22, p99 102, max 3,510. has_single_facility is true on 393,714 studies; has_us_facility on 194,191.
ClinicalTrials.gov protocol elements record participating locations as facilities — name, city, state, country, ZIP — not as a free-text “number of countries” workload field [1]. AACT splits that information. The countries table is one row per country associated with a study, with a removed flag. calculated_values stores number_of_facilities, has_single_facility and has_us_facility [2][13]. Those are different grains. Mixing them is the multicenter error this section exists to stop.
AACT countries in the 1 August 2026 files: 808,912 rows, of which 35,778 are flagged removed. Per-study country counts in this paper use non-removed rows only. Joined to 595,630 NCTs:
| Statistic | Countries per study (non-removed) | Facilities per study |
|---|---|---|
| n | 595,630 | 595,630 |
| Minimum | 0 | 0 |
| Median (p50) | 1 | 1 |
| p75 | 1 | 1 |
| p90 | 1 | 8 |
| p95 | 2 | 22 |
| p99 | 12 | 102 |
| Maximum | 59 | 3,510 |
| Mean | 1.3 | 5.87 |
ClinicalTrials.gov / AACT — EClinCloud analysis, accessed August 2026 [2][3].
Countries and facilities per study are different grains
Both medians are 1. At p90, countries are still 1 while facilities are 8; at p99, 12 countries versus 102 facilities. Equating country rows with investigational sites understates one-country multicenter studies and mis-sizes two-country studies.
The two distributions share a median of 1 and then diverge. At p90, countries are still 1 and facilities are 8. At p95, countries are 2 and facilities are 22. At p99, countries are 12 and facilities are 102. At the maximum, 59 countries versus 3,510 facilities. A country is a jurisdiction in the public file. A facility is a listed location. One country can hold hundreds of facilities. Using “countries” as a synonym for “sites” understates the US-only or China-only multicenter study and overstates a two-country study with one site each.
Multi-country studies — two or more non-removed countries — are 42,756. That is 7.2% of 595,630. Ninety-odd percent of the file is zero or one country. The multinational EDC conversation is about a thin slice of registered studies, and about a thick slice of industry Phase 3, which is the next section.
Share of ClinicalTrials.gov studies with two or more countries
42,756 of 595,630 studies (7.2%) list two or more non-removed countries. The multinational EDC conversation is a thin slice of the register and a thick slice of industry Phase 3.
Facility flags from AACT calculated values, counted on the 595,630-NCT join:
has_single_facilitytrue: 393,714 (66.1%)has_us_facilitytrue: 194,191 (32.6%)
number_of_facilities can be 0 when sites are not listed. A zero in that column is missing location data, not proof of a virtual trial. The single-facility flag is the closer public proxy for “this looks like one listed site,” and even then it is not a CRF-page count.
Top country rows — not distinct studies — among non-removed country records:
| Country name | Country rows |
|---|---|
| United States | 194,045 |
| China | 52,726 |
| France | 42,588 |
| Canada | 32,278 |
| United Kingdom | 28,653 |
| Germany | 28,378 |
| Italy | 25,913 |
| Turkey (Türkiye) | 25,433 |
| Spain | 25,295 |
| South Korea | 17,552 |
AACT countries (non-removed), 1 August 2026 — EClinCloud analysis, accessed August 2026 [2].
United States 194,045 country rows sit next to has_us_facility true on 194,191 studies. The numbers are close because they are related objects, and they are not identical: one is a country-table row, the other is a calculated flag. China 52,726 is country rows. A multinational study that lists China contributes one China row and also contributes rows for every other listed country. Ranking these as “studies per country” double-counts every multinational NCT. They are a geography of rows.
Removed rows are not a rounding error. 35,778 country rows carry the removed flag. A country that was listed and then removed is still in the table. Counting it as a live participating country invents a jurisdiction the current record no longer claims. This paper excludes them from per-study country counts. Any downstream dashboard that skips the flag will inflate multinational share.
Action after this section: never equate countries.name with investigational sites. If you need a public proxy for site count, use facility counts when they are present, then still ask the protocol for the live site list. If the NCT is in the 7.2% with two or more countries, the next question is not “how many countries” as a trophy. It is which languages, which ethics pathways, and which country-level workflows the EDC must actually carry — none of which the country row stores.
The Phase 3 industry interventional cohort
Takeaway: The table a build lead actually needs is not the all-study median. On 24,676 Phase 3 industry interventional studies, outcomes median 6 / p90 20 / p99 52 / max 214 / mean 9.4, and countries median 1 / p75 7 / p90 17 / mean 5.45 / max 59. Observational studies (n = 139,061) have median 3 outcomes.
Filter, stated so it can be reproduced: ClinicalTrials.gov phase string contains PHASE3 (including Phase 2/3 combinations such as PHASE2;PHASE3), sponsor class INDUSTRY, study type INTERVENTIONAL. That cut is 24,676 studies — 4.1% of the 595,630-row file [2][3]. It is not “all Phase 3.” A clean PHASE3 label alone is 41,964 studies, of every sponsor class and including non-interventional misreads if any exist; our cohort is narrower and also slightly wider than a pure PHASE3 token because it includes combination strings.
| Statistic | Outcomes (P3 industry interventional) | Countries (P3 industry interventional) | Outcomes (observational) |
|---|---|---|---|
| n | 24,676 | 24,676 | 139,061 |
| Minimum | 0 | 0 | 0 |
| Median (p50) | 6 | 1 | 3 |
| p75 | 12 | 7 | 5 |
| p90 | 20 | 17 | 10 |
| p95 | 28 | 24 | 14 |
| p99 | 52 | 35 | 29 |
| Maximum | 214 | 59 | 2,334 |
| Mean | 9.4 | 5.45 | 4.5 |
ClinicalTrials.gov / AACT — EClinCloud analysis, accessed August 2026 [2][3].
Countries per study: all ClinicalTrials.gov versus Phase 3 industry interventional
Even in the 24,676-study Phase 3 industry interventional cohort the median country count is still 1. The multinational surface is the upper half: p75 = 7, p90 = 17, mean 5.45. The all-study p90 remains 1. Quote this cohort with its own percentiles.
Two facts in that table should sit on a sticky note.
First: the median country count is still 1, even in this industry Phase 3 interventional cohort. A Phase 3 industry study is not automatically multinational. Half of this cohort lists one country or none. The multinational surface lives in the upper half: p75 = 7 countries, p90 = 17, p95 = 24, p99 = 35, max 59. Mean countries 5.45 versus the all-study mean 1.3 is the tail. Quote p75 and p90 for this cohort, not the all-study p90 of 1.
Second: outcomes move, but less dramatically than countries at the top. Median 6 versus 4 in the all-study file; p90 20 versus 13; p99 52 versus 36; max 214 versus 2,334. The all-study maximum is not in this cohort. It is in observational (max 2,334, median 3, mean 4.5). Using the all-study max to scare a Phase 3 industry quote is looking at the wrong row. Using the observational median to staff a p90 Phase 3 industry study (20 outcomes, 17 countries) is looking at the wrong cohort.
The before-and-after for a build lead:
- Before: “It’s Phase 3, so use the big template.” After: read whether this NCT is in the lower half of the industry Phase 3 country distribution (median 1) or in the p75–p90 band (7–17 countries). Those are different translation, ethics and UAT objects.
- Before: “Observational is always light.” After: observational median is light (3 outcomes). Observational maximum is the heaviest outcome row in the entire file. Band the NCT. Do not band the study type alone.
- Before: one complexity score, maybe a vendor hour percentage. After: a percentile table for the cohort that matches the protocol.
Primary purpose TREATMENT (289,414 design rows) and allocation RANDOMIZED (299,465) describe the file’s common interventional pattern. They still do not replace the Phase 3 × industry × interventional cut. A randomized two-arm academic Phase 2 in one country is not this cohort. A 17-country industry Phase 3 is. The public file can tell you which one you have, once you filter.
Company book-of-business figures are not a denominator here. EClinCloud publishes a 1,100+ clinical-trial claim on its site; that number is a company claim, not a census against which these 24,676 studies can be rated, and it is not used in any percentage in this paper [14].
CTIS contrast, same warning
Takeaway: The CTIS public file is 12,123 EU medicine trials with a country list filled on all of them: median 1 country, p90 8, mean 3.18, max 25. It is a different grain from ClinicalTrials.gov country rows (median 1, p90 1, mean 1.3, max 59). Do not average the two files. Do not retell identity-prefix joins.
Regulation (EU) No 536/2014 applies to clinical trials of medicinal products in the Union. It does not apply to non-interventional studies [15]. CTIS is the information system for that regulation. The public portal is a searchable website of EU/EEA medicine trials from 31 January 2022, under the transparency rules that apply from June 2024 [16][17]. That legal perimeter is why the file cannot be dropped onto ClinicalTrials.gov as “the European subset.”
The 30 July 2026 public CTIS summary extract used here holds 12,123 trials. Country lists are filled on all 12,123. Countries per trial:
| Statistic | CTIS public trials | AACT / ClinicalTrials.gov studies |
|---|---|---|
| n | 12,123 | 595,630 |
| Minimum | 1 | 0 |
| Median (p50) | 1 | 1 |
| p75 | 4 | 1 |
| p90 | 8 | 1 |
| p95 | 11 | 2 |
| p99 | 16 | 12 |
| Maximum | 25 | 59 |
| Mean | 3.18 | 1.3 |
CTIS public summaries, 30 July 2026, and ClinicalTrials.gov / AACT — EClinCloud analysis, accessed August 2026 [2][16]. CTIS trialCountries is an EU-trial country list. AACT countries are ClinicalTrials.gov country rows. Do not rank the two files as the same grain.
Country multiplicity: CTIS public trials versus ClinicalTrials.gov
Medians are both 1. Means (3.18 vs 1.3) and p90s (8 vs 1) are not a ranking of 'who is more multinational.' CTIS is 12,123 public EU medicine trials with a country list on every record; AACT counts ClinicalTrials.gov country rows, including studies with none listed. Do not average the files.
Both medians are 1. The means are not: 3.18 versus 1.3. The p90s are not: 8 versus 1. A reader who subtracts those means and concludes “EU trials are 2.4 countries more multinational” has averaged two objects. CTIS is medicine trials under 536/2014, public in this portal, with a country token on every record (minimum 1). ClinicalTrials.gov is interventional, observational, expanded access and blank, global, with a country minimum of 0 because location data can be missing. The CTIS maximum of 25 is not a contradiction of the AACT maximum of 59. It is a different file, a different legal object, a different tail.
Country tokens in the CTIS extract include a status suffix (Germany:8). This paper counts the length of the list as multiplicity. It does not interpret the numeric suffix as a site count or as a CT.gov-style status. Misreading the suffix as “eight German sites” would manufacture a facility grain the token does not contain.
What this contrast is for: a second-registry check that multinational share is file-dependent. What it is not for: a ranking of “who runs more multinational trials,” a join key to NCT, or a retelling of prefix stripping. EClinCloud’s 24 August 2026 Deep Research paper is the identity workpaper — NCT, CT number, CTR, ICTRP TrialID, and why a failed join is not a non-registration finding [18]. This paper stops at multiplicity. If you need to know whether the NCT and the CT number are the same study, go there. If you need to know how many countries the public object lists, stay here and keep the grains named.
Action after this section: when a protocol is an EU medicine trial, read CTIS country-list length as a CTIS fact. When it is an NCT, read AACT country counts as an NCT fact. If both identifiers exist, you have two public objects. Averaging them is not a reconciliation.
How to quote a build without a fake hour claim
Takeaway: Band the NCT on public multipliers, then list what the file cannot see — visits, CRF pages, edit checks, translations. Do not convert a percentile into a percentage time saving. EDC study-build is the job that remains after the band is named; RTSM and CTMS are different jobs.
The public file can support a band. It cannot support an hour. The honest quote sequence is:
- Read study type, sponsor class and phase from the NCT [3].
- Count specified outcomes, design groups, non-removed countries and listed facilities [2].
- Place the NCT on the matching distribution — all-study, observational, or Phase 3 industry interventional — using median and p90, not the mean.
- Write down the objects the file does not contain.
- Size visits, forms, edit checks, languages and integrations from the protocol and the EDC specification, not from the registry row.
A working band table, using only this snapshot:
| Public-file band | What the file typically shows | What to lock before quoting | What you still cannot see |
|---|---|---|---|
| Observational, body of the file | Median 3 specified outcomes; usually one country | Study type is actually observational; outcome list is complete | Visits, forms, edit checks, source-data flow |
| Median NCT | 4 outcomes, 2 groups, 1 country, 1 facility | You are not silently in the Phase 3 industry tail | Same, plus whether facilities were even listed |
| Phase 3 industry interventional, lower half | Median 6 outcomes; median 1 country | Industry × interventional × Phase 3, and country count is still 1 | Translations may still exist inside one country |
| Phase 3 industry interventional, p75–p90 | ~12–20 outcomes; 7–17 countries | Language matrix, country workflows, UAT breadth | Hours; CRF pages; edit-check count |
| Tail | p99 outcomes 36 (52 in the P3 industry cohort); p99 facilities 102; max outcomes 2,334 is observational | Do not let max or mean set the rate card | Anything operational beyond the row |
ClinicalTrials.gov / AACT — EClinCloud analysis, accessed August 2026 [2][3].
The file cannot see, at minimum:
- Visit schedule. Specified outcomes do not enumerate visits. A four-outcome protocol can have twenty visits; a twenty-outcome protocol can collect everything at three time points.
- CRF pages and fields. One outcome can be one item or a battery.
- Edit checks. Cross-form, cross-visit and query logic are not registry data elements [1].
- Translations and linguistic validation. Country count is not language count. One country can be multilingual; two countries can share a language.
- Integrations. Labs, eCOA, RTSM, imaging, safety. Adjacent work on COA language in public rows is a separate paper [10].
- Actual enrollment. The enrollment field is a submitted number, often anticipated; 7,125 studies in this ClinicalTrials.gov snapshot have a blank enrollment value [3]. It is not subjects in the database.
- Hours. No public column stores study-build effort.
A vendor claim of the form “cuts build time by X%” cannot be checked against this file, and this paper does not make one. Percentile bands are not percentages of time.
EDC is the system that holds the visits, forms and checks once they exist. EClinCloud’s EDC is built around study build as the data manager’s job; study-build and configuration services sit between the protocol and a running database — protocol review, visit and CRF design, edit checks, validation, language setup [14][19]. That is a soft handoff after the band is named, not a substitute for the band.
Keep the products distinct. Randomization and trial supply are RTSM [11]. Portfolio, site and monitoring workflows are CTMS [12]. A 17-country industry Phase 3 will usually need more than one of these. Counting countries in AACT does not tell you which. It tells you the EDC country surface is not the median NCT’s.
Action after this section: put the NCT in a band, list the unseen objects, and quote the build from the protocol. If you need a configured database, that is EDC study-build. If you need kit logic, that is RTSM. If you need site milestones, that is CTMS. If someone offers a universal hour saving derived from “protocol complexity,” ask which distribution, which percentile, and which objects were actually counted.
Frequently asked questions
Takeaway: The 2,334-outcome row is real and observational; country is not facility; CTIS is not ClinicalTrials.gov; “complexity is rising” is not this paper’s claim; enrollment is not enrolled.
Is the 2,334-outcome study a data error?
It is a real row in the 1 August 2026 AACT join: specified-outcome maximum 2,334, and the observational-cohort maximum is the same number [2]. Measure text is blank on only 9 of 3,716,648 outcome rows, so the tail is not a blank-field artefact. It is still an outlier. Median 4, p99 36, mean 6.23 in the all-study file; Phase 3 industry interventional maximum 214. Do not let 2,334 set a mean-based quote. Do not discard it as impossible without opening the NCT. Band it as tail, then read the protocol.
If a study lists two countries, is that two sites?
No. Countries median 1, p90 1, p95 2; facilities median 1, p90 8, p95 22, p99 102, max 3,510 [2]. A country row is a jurisdiction on the AACT countries table (non-removed). A facility is a listed location. Two countries can be two sites or two hundred. One country can be 3,510 listed facilities. has_single_facility is true on 393,714 studies; that flag is the closer public proxy for one listed site, and it is still not a CRF count. Facility counts can be 0 when sites were not submitted.
Why is only 7.2% of ClinicalTrials.gov multi-country if so many trials feel global?
Because the file is mostly not industry Phase 3. 42,756 of 595,630 studies have two or more non-removed countries [2][3]. OTHER sponsors are 71.5% of the file; observational studies are 23.3%. In the Phase 3 industry interventional cohort the median country count is still 1, but p75 is 7 and p90 is 17. Global operational experience is concentrated in a cohort that is 4.1% of registered studies. The all-study percentage answers a different question than the protocol on an industry Phase 3 desk.
Can I average CTIS and ClinicalTrials.gov country counts?
No. CTIS public summaries in this snapshot: 12,123 EU medicine trials, country list filled on all, median 1, p90 8, mean 3.18, max 25 [16]. ClinicalTrials.gov / AACT: 595,630 studies, median 1, p90 1, mean 1.3, max 59 [2]. Regulation (EU) No 536/2014 is medicine trials, not non-interventional studies [15]. CTIS tokens can carry a status suffix; this paper counts list length. Identity joins — NCT against CT number, prefix stripping, ICTRP — belong to the 24 August 2026 paper [18].
Are protocols getting more complex?
This paper does not compute a time trend. It computes a sizing table. The literature slogan is thin relative to the file: in a 101,685-record Europe PMC eClinical union, title/abstract hits for protocol complexity / case report form / eCRF were 74 [20]. NIH RePORTER’s 31-phrase eClinical corpus (4,876 unique projects) retrieves 3 projects for the exact phrase “electronic case report form,” versus 1,023 for “electronic data capture” [21]. That is a phrase boundary. It is not evidence that nobody uses eCRFs, and 74 literature hits are not a workload model. Use the distributions.
Does a blank enrollment field mean the study enrolled nobody?
No. Enrollment on ClinicalTrials.gov is a submitted data element, often anticipated, not a count of subjects in the EDC [1]. 7,125 studies in the 25 July 2026 file have blank enrollment [3]. A filled number is still not “actual enrolled.” Do not size database volume or site load from that field.
Does masking NONE mean the build is easier?
No. NONE is 249,894 of 591,898 design rows [2]. Open-label is common and can coincide with heavy safety, imaging review and oncology endpoint workflows. Masking records the blinding description. It does not record ethics committee load or independent review.
Can I convert specified outcomes into CRF pages or hours?
No. 3,716,648 outcome rows are names and types (primary 1.17 million, secondary 2.34 million) [2]. They are not pages, fields, edit checks or hours. Convert the count into a band, then design the CRF from the protocol.
Methodology and limitations
Takeaway: Three evidence lanes, named grains, dated snapshots. ClinicalTrials.gov and AACT are one register. CTIS is a second. Europe PMC and NIH RePORTER are a slogan check, not a workload model. Filters have edges; we report them.
- Lane 1 — ClinicalTrials.gov / AACT (one register). Study file: ClinicalTrials.gov publicly downloadable records, 25 July 2026, 595,630 unique NCT rows [3][5]. Relational protocol objects: AACT pipe-delimited files, Clinical Trials Transformation Initiative, 1 August 2026 —
calculated_values(596,685 rows),design_outcomes(3,716,648),design_groups(1,097,076),countries(808,912 rows; 35,778 removed),designs(591,898) [2][13][22]. Per-study distributions join AACT objects to the 595,630-NCT file. The August AACT versus July study-file gap is why calculated_values (596,685) and designs (591,898) do not equal 595,630. Protocol data-element definitions are NLM’s, not ours [1][4][7]. Computed 23 August 2026. ClinicalTrials.gov / AACT — EClinCloud analysis, accessed August 2026. - Phase 3 industry interventional cohort. Phase string contains
PHASE3(including PHASE2;PHASE3 combinations), sponsor class INDUSTRY, study type INTERVENTIONAL. n = 24,676. Observational cohort: study type OBSERVATIONAL, n = 139,061, matching the study file. - Country rules. Per-study country counts exclude AACT rows with
removedtrue. Multi-country = two or more non-removed countries. Top-country table is row counts, not distinct studies.has_us_facilityandhas_single_facilityare AACT calculated flags, not the countries table. - Lane 2 — CTIS public summaries. 30 July 2026 extract, 12,123 trials.
trialCountrieslist length is multiplicity; status suffixes on tokens are not site counts. Grain: EU medicine trials under Regulation (EU) No 536/2014, public portal [16][17][15]. Not averaged with ClinicalTrials.gov. Identity joins are out of scope [18]. - Lane 3 — slogan check. Europe PMC eClinical union, 1 August 2026, 101,685 records; title/abstract regex for protocol complexity / case report form / eCRF = 74 hits [20]. NIH RePORTER 31-phrase eClinical corpus, 23 August 2026, 4,876 unique projects; query
"electronic case report form"retrieved 3;"electronic data capture"retrieved 1,023 [21]. Phrase volume is not adoption and not a CRF census. - Rejected for this question. China Drug Trials list (no visit/form/outcome counts). ICTRP (identity, not multiplicity). Full facility address lists (not needed; calculated_values already holds facility counts). Duplicate copies of the same ClinicalTrials.gov, AACT, CTIS and Europe PMC sources inspected and not counted as extra lanes.
- Limitations. Specified outcome count ≠ CRF pages, edit checks or visits. Facility count can be 0 when sites are unlisted. Enrollment ≠ actual enrolled. Max 2,334 is a real observational row and must not set the mean. CTIS and AACT country means are not comparable grains. Masking NONE ≠ simpler ethics. Group-type and country-name tables are rows. This paper does not estimate hours and does not claim a time-saving percentage.
- Client claims. EDC, RTSM and CTMS are kept distinct [14][11][12]. Company trial-volume figures are not denominators. No regulatory-acceptance, enrollment-speed or inspection-readiness claim.
Conclusion
Takeaway: Size the NCT on public multipliers, on the cohort it actually belongs to, without converting a percentile into an hour.
Four takeaways survive the join.
The median NCT is small; the tail is the operational risk. Median 4 specified outcomes and 1 country on 595,630 studies; p90 13 outcomes; 7.2% multi-country. Half of the register is not a 17-country database. Budget the tail separately [2][3].
One NCT is not one workload. INTERVENTIONAL 454,529 versus OBSERVATIONAL 139,061; INDUSTRY 131,151 versus OTHER 425,599. The Phase 3 industry interventional cohort (n = 24,676) has p90 = 20 outcomes and 17 countries, with country median still 1. Observational median outcomes = 3; observational max = 2,334. Filter first [3].
Countries are not sites, and CTIS is not ClinicalTrials.gov. Facilities p90 = 8 when countries p90 = 1. 35,778 country rows are flagged removed. CTIS (12,123 public EU medicine trials, country mean 3.18, p90 = 8) is a 536/2014 object. Do not average it with AACT [16][15].
Quote a band, not a fake hour. The file does not store visits, edit checks, translations or effort. EClinCloud EDC and study-build services exist for the work that starts after the band is named [14][19]. RTSM and CTMS remain separate products [11][12]. No build-time percentage follows from these tables.
If the decision on your desk is how to staff CRF writers and UAT for this NCT, pull the outcome, group, country and facility counts, place them on the matching percentile table, and design from the protocol. The public row is a multiplier. It is not a timesheet.
Sources
1. U.S. National Library of Medicine, Protocol Registration Data Element Definitions for Interventional and Observational Studies — outcome measures, arms/groups, facility location (including country), enrollment, masking, allocation and related protocol elements.
2. Clinical Trials Transformation Initiative, AACT Database — publicly available relational copy of ClinicalTrials.gov protocol and results elements. EClinCloud analysis of the 1 August 2026 pipe-delimited files (design_outcomes, design_groups, countries, designs, calculated_values) joined to the 25 July 2026 study file, accessed August 2026.
3. U.S. National Library of Medicine, ClinicalTrials.gov — publicly downloadable study records. EClinCloud analysis of the 25 July 2026 snapshot (595,630 unique NCT rows: study type, sponsor class, phase, status, enrollment fill, results-posted flag), accessed August 2026.
4. U.S. National Library of Medicine, ClinicalTrials.gov Glossary — study record, NCT number, outcome measure, primary and secondary outcome measure, parallel assignment.
5. U.S. National Library of Medicine, About ClinicalTrials.gov — registry purpose and public access to protocol registration information.
6. U.S. National Library of Medicine, Submit Studies — Protocol Registration and Results System (PRS), the sponsor-facing path for the protocol data elements analysed here.
7. U.S. National Library of Medicine, Results Data Element Definitions for Interventional and Observational Studies — posted results (participant flow, baseline characteristics, outcome measures, adverse events) as a different object from protocol-specified outcomes.
8. International Council for Harmonisation, E6(R3) Guideline for Good Clinical Practice — computerized systems should be proportionate to the trial; cited as a principle, not as a retelling of the guideline.
9. EClinCloud, ICH E6(R3): Clinical trials enter the era of risk-driven digital governance, 10 March 2026 — existing Insight on digital GCP; not this paper’s thesis.
10. EClinCloud, From regulatory intent to practice: choosing a fit-for-purpose clinical outcome assessment, 30 October 2025, and the 28 August 2026 Deep Research article on COA language in public trial records — outcome type taxonomy. This paper uses primary/secondary row counts only as a naming mix.
11. EClinCloud, RTSM — Randomization and Trial Supply Management — distinct from EDC; randomization and supply are not inferred from AACT allocation counts.
12. EClinCloud, CTMS — Clinical Trial Management — distinct from EDC; site and milestone workflows are not inferred from facility counts.
13. Clinical Trials Transformation Initiative, AACT data dictionary / documentation — table and column definitions, including calculated facility flags and country removed.
14. EClinCloud, EDC — Electronic Data Capture — study-build product surface. Company trial-volume and time-saving figures on the site are company claims and are not denominators or findings in this paper.
15. European Union, Regulation (EU) No 536/2014 on clinical trials on medicinal products for human use — Article 1: applies to clinical trials in the Union; does not apply to non-interventional studies.
16. European Medicines Agency, Search clinical trials and reports (CTIS public portal) — public information on EU/EEA clinical trials submitted through CTIS. EClinCloud analysis of 12,123 public summary records (30 July 2026 snapshot; trialCountries list length), accessed August 2026.
17. European Medicines Agency, Clinical Trials Information System — CTIS as the system under the Clinical Trials Regulation; public website and revised transparency rules applicable from June 2024.
18. EClinCloud, Four public trial registers, four identity contracts, 24 August 2026 — identity joins (NCT, CT number, CTR, ICTRP). This paper does not retell prefix or join-key findings.
19. EClinCloud, Professional Services — Study Build & Configuration — protocol review, database and visit design, CRF and edit-check configuration, validation and language setup.
20. Europe PMC — bibliographic database of life-science literature. EClinCloud analysis of a 101,685-record eClinical union (1 August 2026 snapshot): 74 title/abstract hits for protocol complexity / case report form / eCRF, accessed August 2026.
21. U.S. National Institutes of Health, RePORTER — funded-project search. EClinCloud analysis of a 31-phrase eClinical corpus (4,876 unique project applications, 23 August 2026 snapshot): "electronic case report form" retrieved 3; "electronic data capture" retrieved 1,023, accessed August 2026.
22. Clinical Trials Transformation Initiative, CTTI — public-private partnership that maintains AACT as a research copy of ClinicalTrials.gov.