October Data Sharing and Reuse Seminar

Friday, October 9, 2026

Hyunghoon (Hoon) Cho, Ph.D. will present "Enabling Collaborative Genomic Studies with Privacy" from 12:00 p.m.–1:00 p.m. EDT.

About the Seminar

The sensitive nature of genomic data poses major challenges for data sharing and collaboration in biomedicine. Traditional safeguards often lead to data silos, hindering large-scale analysis. I will describe our recent work on secure federated (SF) algorithms, which combine cryptography and distributed computation to enable collaborative genomics research without compromising privacy. I will showcase practical tools we have developed for key tasks, including genome-wide association studies (Nature Genetics, 2025) and inference of genetic relatedness (Genome Research, 2024). Finally, I will discuss our ongoing efforts to deploy these methods across large-scale biobanks and highlight broader opportunities for privacy-enhancing technologies in biomedical data sharing.

About the Speaker

Hyunghoon (Hoon) Cho, PhD, is an Assistant Professor of Biomedical Informatics & Data Science and of Computer Science at Yale University. He received his Ph.D. in Electrical Engineering and Computer Science at MIT in 2019. Previously, he received his M.S. and B.S. with Honors in Computer Science from Stanford University. His research focuses on overcoming key computational challenges in analyzing massive and distributed biomedical data, developing modern tools based on applied cryptography and machine learning. He is especially interested in building privacy-preserving, scalable AI and mathematical models to unlock insights from complex genomic data. He is a recipient of the NIH Director's Early Independence Award.

About the Seminar Series

The seminar is open to the public and registration is required each month. Individuals who need interpreting services and/or other reasonable accommodations to participate in this event should contact Allison Hurst at 301-670-4990. Requests should be made at least five days in advance of the event. 

The National Institutes of Health (NIH) Office of Data Science Strategy hosts this seminar series to highlight examples of data sharing and reuse on the second Friday of each month at noon ET. The monthly series highlights researchers who have taken existing data and found clever ways to reuse the data or generate new findings. A different NIH institute or center will also share its data science activities each month.

September Data Sharing and Reuse Seminar

Friday, September 11, 2026

Rui Feng, Ph.D. will present "Sex Differences in Acute Kidney Injury Risk and Outcomes: Insights from MIMIC-III and MIMIC-IV" from 12:00 p.m.–1:00 p.m. EDT.

About the Seminar

Acute kidney injury (AKI) affects 13-18% of U.S. hospitalizations and increases the risks of chronic kidney disease (CKD) and death. However, which clinical factors are associated with AKI and its sequelae, and whether these associations differ between women and men, remain poorly understood. This study combines machine learning and classical regression models to examine predictors across the continuum from before ICU admission, for AKI onset, post-AKI CKD and death.

We analyzed adult ICU admissions from MIMIC-III (2001–2012) and MIMIC-IV (2008–2019). Six machine-learning models were developed in MIMIC-IV and validated in MIMIC-III using demographic, clinical, laboratory, treatment, and clinical-text features.

Performance was assessed through discrimination, diagnostic accuracy, and calibration. Sex differences in predictor associations were evaluated using regression interaction terms and Friedman’s H-statistics. Progression from AKI to CKD and death was evaluated using competing-risk regression and DeepSurv.

XGBoost achieved the strongest discrimination for AKI, and DeepSurv performed best in predicting CKD and death. Creatinine, bicarbonate, and hemoglobin showed differential associations with AKI risk between women and men. Sex differences involving baseline creatinine and potassium were also observed for CKD and mortality after AKI. These findings help explain heterogeneity in AKI risk and progression and may inform the development and evaluation of more individualized risk-assessment and surveillance strategies.

The talk will also illustrate how human researchers and AI collaborated throughout the project, from refining research questions to evaluating models and interpreting results.

About the Speaker

Rui Feng, Ph.D., is an Associate Professor in the Department of Biostatistics, Epidemiology, and Informatics at the University of Pennsylvania Perelman School of Medicine. Her research focuses on statistical and machine-learning methods for high-dimensional biomedical data, with particular interests in deep learning, explainable AI, and modeling complex interactions that influence health outcomes.

About the Seminar Series

The seminar is open to the public and registration is required each month. Individuals who need interpreting services and/or other reasonable accommodations to participate in this event should contact Allison Hurst at 301-670-4990. Requests should be made at least five days in advance of the event. 

The National Institutes of Health (NIH) Office of Data Science Strategy hosts this seminar series to highlight examples of data sharing and reuse on the second Friday of each month at noon ET. The monthly series highlights researchers who have taken existing data and found clever ways to reuse the data or generate new findings. A different NIH institute or center will also share its data science activities each month.

Catalyzing Biomedical Data Innovation: Three Years of the DataWorks! Prize

Wednesday, August 5, 2026

Welcome to the August 2026 Director's Corner!

At the NIH Office of Data Science Strategy (ODSS), our mission is to build a cutting-edge biomedical data ecosystem - one that is fully integrated and rooted in FAIR principles (Findable, Accessible, Interoperable, and Reusable). In this post, I want to celebrate a program that has become a powerful engine for that mission: the DataWorks! Prize.


A Competition That Keeps Raising the Bar

Launched in partnership with the Federation of American Societies for Experimental Biology (FASEB), the DataWorks! Prize recognizes researchers who demonstrate the real-world impact of good data management and sharing practices. Over three years, it has grown from honoring foundational data sharing into a dynamic competition spotlighting cutting-edge secondary analysis - and it is driving meaningful change across the research community. 

Here's a look at how the competition evolved.


Year 1: Foundational Data Sharing and Reuse (2022)

DataWorks! Prize year one image

The inaugural competition set the stage by recognizing teams that demonstrated the transformative power of data sharing and high-quality data resources to advance human health. With 106 registered teams and 2,704 team members, the response was extraordinary.

Standout winners included the BrainChart team, whose work on human lifespan brain charts opened new frontiers in neuroscience, and the National COVID Cohort Collaborative (now evolved into the National Clinical Cohort Collaborative), which broke down barriers to clinical data access at a critical moment in public health. This first year also helped align the research community with the then-upcoming NIH Data Management and Sharing Policy- building good habits from the ground up.


Year 2: Best Practice "Recipes" (2023)

The 2023 competition pivoted toward reproducibility - showcasing methodologies so clearly documented that other researchers could pick them up and run with them. These "recipes" spanned diverse disciplines including genomics, immunology, and neuroscience, drawing 39 teams and 220 researchers.

A highlight was the COVID-19 and Cancer Consortium (CCC19), recognized for its remarkable collaborative efforts in collecting and sharing critical patient data across institutions. The message was clear: great science isn't just about findings - it's about making those findings repeatable and reusable.


Year 3: Data Reuse from Generalist Repositories (2024)

The most recent competition raised the stakes further, focusing on innovative secondary analysis using data from generalist repositories participating in the Generalist Repository Ecosystem Initiative (GREI) program. The 2024 competition attracted 82 teams and 311 team members spanning biochemistry, clinical research, genomics, immunology, molecular biology, and neuroscience.

The Grand Prize went to Divergent Equilibria for their groundbreaking work on sex differences in Acute Kidney Injury — a compelling example of what's possible when high-quality shared data meets creative scientific inquiry.


Three Years, Over One Million Dollars, Countless Discoveries

Across its three years, the DataWorks! Prize has awarded over $1 million to teams exemplifying excellence in data science. But the impact goes beyond the prize money. Winning teams are invited to share their experiences at NIH forums, turning their work into blueprints for researchers around the world.

They are more than competition winners - they are champions of a new culture in biomedical research, one where data is treated as a shared resource and a springboard for discovery.

Thanks also to our partners in HeroX for administering the competition. For more information on the DataWorks! Prize and past winners, visit Data Science at NIH.


Stay connected with ODSS at our website and subscribe to our newsletter for the latest news as we continue turning data into groundbreaking discoveries.

August Data Sharing and Reuse Seminar

Friday, August 14, 2026

Stephen C.J. Parker, Ph.D. will present "PanKbase: An Open, FAIR Knowledge Base of the Human Pancreas to Accelerate Diabetes Research" from 12:00 p.m.–1:00 p.m. EDT.

About the Seminar

PanKbase, the Pancreas Knowledge Base, is an open, NIDDK-supported resource that organizes, harmonizes, and disseminates data on the human pancreas to accelerate diabetes research, with an initial focus on type 1 diabetes. Multimodal pancreas data are generated across many programs but remain fragmented and processed in inconsistent ways, which makes them hard to discover and reuse. PanKbase brings these data together into a centralized, computation-ready, FAIR resource spanning genomic, epigenomic, transcriptomic, and physiological measurements from human donors. In this talk we will describe how PanKbase applies FAIR principles and uniform reprocessing to turn disparate datasets into a reusable foundation for the whole community, from basic biomedical research scientists to the machine learning field. I will highlight how open sharing and reuse can generate new insight into pancreatic biology and speed progress toward preventing and treating diabetes.

About the Speaker

Stephen C.J. Parker, Ph.D., is a tenured Professor of Computational Medicine and Bioinformatics at the University of Michigan, with secondary appointments in the Departments of Human Genetics and Biostatistics. He directs the Epigenomic Metabolic Medicine Center and the Integrative Data Analytics Core. His group develops integrative computational methods that interpret non-coding DNA and single-cell multi-omic tissue data to translate diabetes and related metabolic trait genome-wide association signals into biological mechanisms and therapeutic insights. He is a standing member of the NIH Genetics of Health and Disease Study Section and received the American Diabetes Association Outstanding Scientific Achievement Award in 2024. He is a Principal Investigator and Steering Committee Member of PanKbase.

About the Seminar Series

The seminar is open to the public and registration is required each month. Individuals who need interpreting services and/or other reasonable accommodations to participate in this event should contact Allison Hurst at 301-670-4990. Requests should be made at least five days in advance of the event. 

The National Institutes of Health (NIH) Office of Data Science Strategy hosts this seminar series to highlight examples of data sharing and reuse on the second Friday of each month at noon ET. The monthly series highlights researchers who have taken existing data and found clever ways to reuse the data or generate new findings. A different NIH institute or center will also share its data science activities each month.