Creating Citation Reports with InCites
October Data Sharing and Reuse Seminar
Friday, October 9, 2026
Hyunghoon (Hoon) Cho, Ph.D. will present "Enabling Collaborative Genomic Studies with Privacy" from 12:00 p.m.–1:00 p.m. EDT.
About the Seminar
The sensitive nature of genomic data poses major challenges for data sharing and collaboration in biomedicine. Traditional safeguards often lead to data silos, hindering large-scale analysis. I will describe our recent work on secure federated (SF) algorithms, which combine cryptography and distributed computation to enable collaborative genomics research without compromising privacy. I will showcase practical tools we have developed for key tasks, including genome-wide association studies (Nature Genetics, 2025) and inference of genetic relatedness (Genome Research, 2024). Finally, I will discuss our ongoing efforts to deploy these methods across large-scale biobanks and highlight broader opportunities for privacy-enhancing technologies in biomedical data sharing.
About the Speaker
Hyunghoon (Hoon) Cho, PhD, is an Assistant Professor of Biomedical Informatics & Data Science and of Computer Science at Yale University. He received his Ph.D. in Electrical Engineering and Computer Science at MIT in 2019. Previously, he received his M.S. and B.S. with Honors in Computer Science from Stanford University. His research focuses on overcoming key computational challenges in analyzing massive and distributed biomedical data, developing modern tools based on applied cryptography and machine learning. He is especially interested in building privacy-preserving, scalable AI and mathematical models to unlock insights from complex genomic data. He is a recipient of the NIH Director's Early Independence Award.
About the Seminar Series
The seminar is open to the public and registration is required each month. Individuals who need interpreting services and/or other reasonable accommodations to participate in this event should contact Allison Hurst at 301-670-4990. Requests should be made at least five days in advance of the event.
The National Institutes of Health (NIH) Office of Data Science Strategy hosts this seminar series to highlight examples of data sharing and reuse on the second Friday of each month at noon ET. The monthly series highlights researchers who have taken existing data and found clever ways to reuse the data or generate new findings. A different NIH institute or center will also share its data science activities each month.
September Data Sharing and Reuse Seminar
Friday, September 11, 2026
Rui Feng, Ph.D. will present "Sex Differences in Acute Kidney Injury Risk and Outcomes: Insights from MIMIC-III and MIMIC-IV" from 12:00 p.m.–1:00 p.m. EDT.
About the Seminar
Acute kidney injury (AKI) affects 13-18% of U.S. hospitalizations and increases the risks of chronic kidney disease (CKD) and death. However, which clinical factors are associated with AKI and its sequelae, and whether these associations differ between women and men, remain poorly understood. This study combines machine learning and classical regression models to examine predictors across the continuum from before ICU admission, for AKI onset, post-AKI CKD and death.
We analyzed adult ICU admissions from MIMIC-III (2001–2012) and MIMIC-IV (2008–2019). Six machine-learning models were developed in MIMIC-IV and validated in MIMIC-III using demographic, clinical, laboratory, treatment, and clinical-text features.
Performance was assessed through discrimination, diagnostic accuracy, and calibration. Sex differences in predictor associations were evaluated using regression interaction terms and Friedman’s H-statistics. Progression from AKI to CKD and death was evaluated using competing-risk regression and DeepSurv.
XGBoost achieved the strongest discrimination for AKI, and DeepSurv performed best in predicting CKD and death. Creatinine, bicarbonate, and hemoglobin showed differential associations with AKI risk between women and men. Sex differences involving baseline creatinine and potassium were also observed for CKD and mortality after AKI. These findings help explain heterogeneity in AKI risk and progression and may inform the development and evaluation of more individualized risk-assessment and surveillance strategies.
The talk will also illustrate how human researchers and AI collaborated throughout the project, from refining research questions to evaluating models and interpreting results.
About the Speaker
Rui Feng, Ph.D., is an Associate Professor in the Department of Biostatistics, Epidemiology, and Informatics at the University of Pennsylvania Perelman School of Medicine. Her research focuses on statistical and machine-learning methods for high-dimensional biomedical data, with particular interests in deep learning, explainable AI, and modeling complex interactions that influence health outcomes.
About the Seminar Series
The seminar is open to the public and registration is required each month. Individuals who need interpreting services and/or other reasonable accommodations to participate in this event should contact Allison Hurst at 301-670-4990. Requests should be made at least five days in advance of the event.
The National Institutes of Health (NIH) Office of Data Science Strategy hosts this seminar series to highlight examples of data sharing and reuse on the second Friday of each month at noon ET. The monthly series highlights researchers who have taken existing data and found clever ways to reuse the data or generate new findings. A different NIH institute or center will also share its data science activities each month.
Introducing the Generalist Repository Selection Flowchart v2
Friday, February 27, 2026
In July 2024, members of the Generalist Repository Ecosystem Initiative (GREI) published version 1 of the Generalist Repository Selection Flowchart in
Data Sharing Demystified — For the Community, By the Community
Friday, March 27, 2026
Practical insights, researcher success stories, and expert answers from the NIH GREI webinar series.
GREI Year 4: Reflections, Progress, and What Comes Next
Friday, June 5, 2026
As the Generalist Repository Ecosystem Initiative (GREI) concludes its fourth year and moves into its fifth, this update examines the program’s progre
A Survey of GREI Repositories’ Data Packaging Practices and Plans | by The GREI Community
Thursday, June 18, 2026
Data packaging is the practice of organizing dataset components, including metadata, documentation, and data files, into a well-structured container f