The Accelerator is made up of a coordinated group that has translational data expertise that spans Data Science, Data Engineering and Data Governance.

Data Science

The Translational Data Science group partners with research and clinical teams to turn complex scientific questions into scalable, reusable data assets and computational workflows. Our focus is not only on enabling analysis, but on ensuring that outputs are consistent, reproducible, and broadly reusable across studies.

Meet our Data Science team members.

We operate through two integrated capabilities:

  • Translational Analytics - Via data products & cohort analytics, we transform real-world health data into reliable, analysis-ready translational datasets
  • Research Informatics – We aim to enable scalable, reproducible computational workflows and tools

We develop data products, data packages and data analysis tools to enable effective data use and integration between research and clinical care. We also collaborate with our clinical and research partners to develop statistical/ML/NLP models, data analyses, support applications of LLMs in Databricks using clinical data, and data-driven dashboards and visualizations using real-world healthcare and laboratory-generated data.

By integrating these capabilities, this group:

  • Connects cohort definitions, multimodal data assets, and computational workflows
  • Ensures curated data products and research outputs are scalable, reproducible, and reusable
  • Leads cross-cutting projects that require coordinated data and computational approaches

These capabilities are tightly coupled to ensure that research questions are translated into durable data products and workflows, supporting a continuous learning health system.

We collectively support the Translational project feasiblity Data House Call as well as the DaSL Supported Resources Data House Call.

Translational Analytics

We develop curated clinical data products and cohort definitions from real-world clinical and research data. Our work enables consistent, reusable datasets that support a wide range of downstream analyses.

We collaborate with investigators and clinical teams to:

  • Define computable cohorts from real-world data
  • Develop standardized, analysis-ready data products (OMOP)
  • Improve the quality, consistency, and accessibility of clinical data for research
  • Support CAIA related projects being piloted at Fred Hutch

Our focus is on creating reliable and reusable data assets that reduce duplication of effort and support consistent analysis across studies.

Our work focuses on:

  • The development, release and maintenance of the OMOP data model at Fred Hutch and its de-identified version on Databricks.
  • Support for self-service use of Atlas for cohort discovery and definition, de-identified data use and interpretation of the OMOP data model definitions.
  • Facilitated data extracts for complex data projects when self-service resources or identifiable data is required for research project execution.

You can request support for a cohort analysis and/or data extract via this link: https://centernet.fredhutch.org/u/data-science-lab/data-science/clinical-data-for-research.html.

We support the Using Healthcare Data (OMOP) Data House Call.

Research Informatics

We focus on enabling the use of complex biomedical datasets through scalable computing, reproducible workflows, and research software tools.

We support:

  • Scalable computational workflow development using WDL and tools such as PROOF (software maintained by HCI)
  • Effective high-performance and cloud-based computing (including on our local HPC and in Databricks)
  • Advising for data management, data processing and analysis of research data
  • cBioPortal use, including supporting users in uploading their own data, accessing data from other sources such as GENIE and FHCC oncology patients (the FHCC MOD data model).

The team works closely with Fred Hutch IT’s Scientific Computing group and the Accelerator Engineering team to leverage institutional infrastructure, including on-premise clusters, storage, and research applications as well as Databricks computational resources on CARDS.

We support the Research Computing and Data Management.

WILDS WDL Development Program

We offer free, collaborative support to help Fred Hutch researchers develop WDL (Workflow Description Language) workflows for scalable, reproducible computational analyses. In exchange, completed workflows are contributed to the open-source WILDS WDL Library (GitHub) for the broader research community.

Through this program, we:

  • Collaborate with researchers to convert existing scripts, pipelines, or tool combinations into WDL workflows
  • Provide workflows optimized for Fred Hutch HPC and/or cloud environments
  • Publish all workflows to the WILDS WDL Library under open-source licenses
  • Provide documentation on how to run and adapt workflows

What to Expect:

  1. Submit a request using our intake form
  2. Initial meeting – We’ll schedule a Data House Call to discuss your project, assess fit, and scope the work
  3. Development – The WILDS team writes the WDL workflow
  4. Testing & feedback – You test the workflow with your data and provide feedback
  5. Publication – The workflow is added to the WILDS WDL Library with appropriate attribution

This program works best when you have existing scripts or a pipeline you want to scale up, your workflow has broad applicability to other researchers, you’re willing to have the final product shared publicly, and you can commit time to testing and providing feedback.

Ready to get started? Submit a request here.

Data Engineering & Platform

Meet our Data Engineering & Platform team members.
The Data Engineering & Platform team builds and operates the infrastructure, applications, and data pipelines that power the Translational Data Accelerator. The team is responsible for securely integrating, processing, and delivering data while maintaining the reliability, performance, and security of the research data platform. Their work supports applications including Databricks, Atlas, and cBioPortal, providing researchers with scalable access to trusted data and computational resources.

Data Governance

Meet our Data Governance team members.

The Data Governance team enables the responsible use of clinical and research data by developing policies, stewardship practices, and access processes that balance scientific opportunity with privacy, security, and regulatory requirements. The team works closely with researchers and institutional partners to reduce barriers to data access while ensuring that sensitive data are used appropriately, ethically, and in accordance with institutional standards. More about our work here.

Strategy & Coordination

You can find these team members distributed on our main site, including Amy Paguirigan, Sitapriya Moorthi, James Eddy, Maria Owen and Seam Olmstead.

Our Strategy & Coordination staff provide the scientific, technical, and operational leadership that guides the Translational Data Accelerator. Working across data science, engineering, governance, and institutional partners, this staff works to develop strategic priorities, coordinate major initiatives, coordinates relationships across our Cancer Consortium partners, and ensures that platform investments align with the needs of translational research. By bringing together scientific expertise, technical leadership, program administration, and project management, the team helps translate institutional priorities into sustainable capabilities that accelerate discovery.