Building a Study-Ready Database: Protocol to CRF

In a Phase II oncology study, an 80 to 120-page protocol can become 40 to 60 EDC forms, 400 to 800 variables, and 150 to 300 edit checks. That scale gap is a translation problem, not an automatic conversion. Each decision affects whether the database captures the protocol accurately and whether the collected data support the study endpoints.

A clinical trial database is not a general-purpose data repository. It must be built from the protocol, and reading the protocol without a structured extraction method can produce a database shaped by the reader's interpretation instead of the protocol's requirements. That distinction matters during regulatory review, when the configuration must be justified against the approved protocol version.

What "Study-Ready" Means in Practice

A study-ready database gives every protocol collection requirement a corresponding EDC form, field, or derived variable, with edit checks that correctly apply the protocol's constraints. The principle is simple. The work is not. Protocols are narrative regulatory documents, not database specifications, so converting narrative into database structure requires explicit design decisions.

Those decisions include the CDASH domain for each collection type, the field names and labels for each protocol variable, the forms required at each visit, the conditions that enforce protocol range definitions, and the way eligibility criteria become inclusion and exclusion fields. A CDM makes hundreds of such decisions during a study build. Many are routine, while others require sponsor clarification. Without that clarification, assumptions become part of the database.

Reading Visit Schedules Into Form Sequence

The Schedule of Events, also called the Schedule of Assessments, is usually the most directly translatable protocol section. It commonly presents visit names as columns and assessments as rows. That layout maps naturally to an EDC visit and form hierarchy: each column represents a visit event, and each marked cell represents a form assignment.

The table also compresses several layers of conditional logic. Unscheduled visits may have distinct form rules. Some assessments apply only at selected visits within a treatment phase. Early termination visits may resemble end-of-study visits but require modifications. These conditions must be extracted from footnotes and protocol cross-references, not from the table alone.

Visit windows require a separate review. The Schedule of Events identifies the nominal visit day but often omits window definitions, which may appear in a Visit Definition section or in individual narrative sections. The CDM must locate those definitions, compare them with the nominal days in the table, and configure the EDC visit structure accordingly. Differences between the table and narrative should be treated as an expected review point.

Endpoint Definitions to CRF Domains

The Objectives and Endpoints section states what the study is designed to measure. Primary, secondary, and exploratory endpoints each create collection requirements that must be resolved into specific CDASH domains and variables.

Primary efficacy endpoints are often more precise. "Overall Survival defined as time from randomization to death from any cause" requires survival status fields, date of death fields, and censoring logic. Response endpoints such as RECIST, PFS, and ORR require imaging assessment forms with protocol-specific response categories that match the assessment criteria version named by the sponsor.

Secondary and exploratory endpoints are less direct. A secondary endpoint described as "Patient-Reported Outcomes assessed using [instrument name] at scheduled visits" requires the CDM to identify the instrument, obtain its validated item set, apply its scoring guidance, and establish the triggering visits. When the protocol cites an instrument without including its items, the CDM must source them separately. Version-specific scoring requirements must be confirmed before configuration.

Eligibility Criteria to Inclusion/Exclusion Forms

Inclusion and exclusion criteria usually become an I/E Criteria Checklist in the EDC, with each criterion answered at screening. The apparent one-criterion, one-field mapping creates several practical questions.

First, criteria are often compound. "Age 18 to 75 years with confirmed histological diagnosis of [tumor type] and ECOG Performance Status 0 to 2" states three conditions that may need separate fields if the sponsor requires field-level confirmation rather than a single criterion response. The protocol does not dictate how compound criteria should be divided. That database design decision affects data review and eligibility adjudication.

Second, exclusion criteria may refer to laboratory values or findings captured elsewhere. Fields configured for cross-reference with those data, rather than for a simple yes or no response, require additional query logic. The protocol rarely specifies this approach. It generally follows sponsor SOPs or CDM judgment informed by prior study experience.

Where CDASH Guides and Where It Does Not

CDASH, or Clinical Data Acquisition Standards Harmonization, supplies standard variable names, definitions, and domain structures for common clinical data. Applying CDASH conventions during the database build simplifies later SDTM mapping for regulatory submission. It also gives CROs, sponsors, and EDC vendors a shared vocabulary, reducing translation when data move between teams.

CDASH covers core domains well: Demographics (DM), Vital Signs (VS), Laboratory Tests (LB), Adverse Events (AE), Concomitant Medications (CM), Medical History (MH), and Exposure (EX). In these domains, defined names, labels, and value list structures make the database familiar to people who work with CDISC standards.

CDASH offers less direction for study-specific forms, including disease assessments, instrument-based endpoints such as patient-reported outcomes and performance status scales, and protocol-specific derived variables. Here, the CDM must make design choices without a standards template. Those choices should remain documented and defensible: the naming convention, controlled vocabulary, derivation logic, and protocol source for each decision.

CDASH alignment is a map, not proof that the database is correct. A database can use CDASH names accurately while assigning forms to the wrong visits, applying incorrect range logic, or using PRO items that differ from the validated instrument. Standards alignment and protocol alignment are separate verification questions, and both must be answered affirmatively before the database is study-ready.

The Build as Protocol Interpretation Record

Every clinical trial database partly records how the building team interpreted the protocol. When the team documents those interpretations and links each decision to its protocol source, the study build becomes a verification record. When reasoning remains implicit in the CDM's configuration work, the database is harder to audit and update when amendments arrive.

For study teams, build review should therefore ask not only whether the database looks right, but also where each design decision originated. Structured protocol digitization can provide a traceable answer. With manual transcription, the answer may exist only in a CDM's notes, or not at all. The difference becomes especially important when a protocol amendment arrives mid-study and the team must identify exactly which build elements it changes.

Build from protocol, not assumptions

Upload a protocol for a traceable build review, with each database decision linked to its source for sponsor review.