The Translational Data Accelerator develops and maintains trusted data models that make clinical and molecular data easier to discover, analyze, and share for research. These models are built from raw data we ingest from a variety of sources to support the diverse data needs of our translational data community.

Data Models for Translational Research

Research data often originates from many different systems, each designed for operational purposes rather than scientific discovery. Electronic health records, clinical laboratory systems, cancer registries, genomic testing vendors, and research databases frequently use different structures, terminologies, and formats.

Standardized data models help overcome these challenges by organizing information into consistent formats and common vocabularies. By transforming diverse data sources into shared representations centrally, researchers can more easily explore available data, develop reproducible analyses, collaborate across teams and institutions, and leverage a growing ecosystem of community-developed tools.

The Translational Data Accelerator maintains two complementary data models that support translational oncology research at Fred Hutch Cancer Center:

  • The OMOP Common Data Model for standardized clinical and observational research data

  • The Fred Hutch Molecular Oncology Data Model for molecular diagnostics and precision oncology research

OMOP

The OMOP Common Data Model provides a standardized framework for representing observational healthcare data using common structures and vocabularies, enabling consistent analysis across diverse healthcare systems and research organizations. OMOP has become one of the most widely adopted clinical research data standards globally and serves as the foundation for the international OHDSI research network.

The OMOP Common Data Model (OMOP CDM) is the primary standardized clinical research data model used by the Translational Data Accelerator.

OMOP transforms information from electronic health records and related clinical systems into a consistent structure with standardized terminology. Our implementation integrates clinical data from multiple institutional sources to improve consistency and interoperability of the raw data. This harmonization allows researchers to work with data using common definitions for diagnoses, medications, procedures, laboratory results, and clinical events regardless of the original source system.

At Fred Hutch, OMOP serves as the foundation for:

  • Cohort discovery and study feasibility assessments
  • Observational clinical research
  • Clinical trial planning
  • Population-level analyses
  • Reproducible research workflows
  • Data sharing and collaboration with external partners

Researchers can explore de-identified OMOP data through Atlas or request curated identified or de-identified datasets for approved research projects via Databricks.

You can find our release notes for our OMOP data model on our Announcements page.

Fred Hutch Molecular Oncology Data Model

While OMOP provides an excellent foundation for clinical research, many translational oncology studies require detailed molecular information that extends beyond traditional healthcare data standards. The Fred Hutch Molecular Oncology Data Model was developed to support translational oncology research by integrating molecular testing results with standardized clinical data.

This model will combine:

  • Clinical context and patient outcomes derived from OMOP
  • Genomic variant information
  • Copy number alterations
  • Gene expression measurements
  • Protein biomarker results
  • Minimal residual disease (MRD) testing
  • Other specialty laboratory and molecular diagnostic results

The resulting resource provides a unified framework for exploring relationships between molecular characteristics, treatments, and outcomes across patient populations. The model is designed to support both retrospective translational research and emerging precision oncology and clinical trials initiatives.

Researchers can explore de-identified molecular datasets through cBioPortal or request curated identified or de-identified datasets for approved research projects via Databricks.

You can find our release notes for our multimodal data sets available in Databricks and cBioPortal on our Announcements page.

Data Sources

CARDS is meant to aggregate translational data related to the Adult Oncology Program for many downstream uses. As such we regularly ingest new data sources and integrate them into the broader multi-modal data set housed in CARDS.

Current

  • Epic Clarity (the data “behind” Epic)
  • UWM DEEP (legacy and current data in UWM data model, Orca data and legacy EHRs)
  • Caris clinical genomics
  • Tempus clinical genomics
  • Foundation Medicine clinical genomics
  • Guardant Health clinical genomics
  • Fred Hutch Cancer Registry data

Up Next

  • UWLM Clinical laboratory data (Oncoplex, Digital Pathology, other speciality labs)