search

menu

  • Research Research
    • Where science meets inspired minds

    • Back
    • Research
    • Our Science
    • Research Groups
    • Facilities & Platforms
    • Clinical research
    • Find a researcher
    • Publications
    • Knowledge Transfer
  • Careers & study Careers & study
    • Become a leader in cancer research

    • Back
    • Careers & study
    • Vacancies
    • Faculty
    • Scientific staff
    • Scientific support staff
    • Postdoctoral fellows
    • PhD Students
    • Operational staff
    • Clinical fellows
    • Life in Amsterdam
    • Student internships
  • News & Events News & Events
    • Check out our stories and events

    • Back
    • News & Events
    • News
    • Media & Press
    • Calendar
  • About us About us
    • Maximum impact for cancer patients

    • Back
    • About us
    • Our vision
    • Organization
    • Collaborations
    • Responsible Research
    • Support us
    • Visit us
    • Contact us
  • Support us
Support us
  • Home
  • Publications
  • Research
  • Publications
  • Article

Actionability of Synthetic Data in a Heterogeneous and Rare Health Care Demographic: Adolescents and Young Adults With Cancer.

Joshi Hogenboom ,
Aiara Lobo Gomes ,
Andre Dekker ,
Winette Van Der Graaf ,
Olga Husson ,
Leonard Wee

Abstract

METHODS

A population-based cross-sectional cohort study of 3,735 AYAs was subsampled at random to produce 13 training data sets of varying sample sizes. We studied four distinct generator architectures built on the open-source Synthetic Data Vault library. Each architecture was used to generate SD of varying sizes on the basis of each aforementioned training subsets. SD actionability was assessed by comparing the resulting SD with their respective real data against three metrics-veracity, utility, and privacy concealment.

CONCLUSION

SD is a potentially promising option for data sharing and data augmentation, yet sample size plays a significant role in its actionability. SD generation should go hand-in-hand with consistent scrutiny, and sample size should be carefully considered in this process.

RESULTS

All examined generator architectures yielded actionable data when generating SD with sizes similar to the real data. Large SD sample size increased veracity but generally increased privacy risks. Using fewer training participants led to faster convergence in veracity, but partially exacerbated privacy concealment issues.

PURPOSE

Research on rare diseases and atypical health care demographics is often slowed by high interparticipant heterogeneity and overall scarcity of data. Synthetic data (SD) have been proposed as means for data sharing, enlargement, and diversification, by artificially generating real phenomena while obscuring the real patient data. The utility of SD is actively scrutinized in health care research, but the role of sample size for actionability of SD is insufficiently explored. We aim to understand the interplay of actionability and sample size by generating SD sets of varying sizes from gradually diminishing amounts of real individuals' data. We evaluate the actionability of SD in a highly heterogeneous and rare demographic: adolescents and young adults (AYAs) with cancer.

More about this publication

JCO clinical cancer informatics

Volume 8
Pages e2400056
Publication date 01-12-2024

Full text links

Publisher website (DOI) 10.1200/CCI.24.00056
Europe PubMed Central 39626135
Pubmed 39626135

Where science meets inspired minds

Contact

Plesmanlaan 121
1066CX Amsterdam

020 512 9111 communicatie@nki.nl

Quick links

  • Vacancies
  • News
  • Contact us
  • Media & Press

Follow us on

Disclaimer
Privacy statement
Cookies
Change cookie settings

This site uses cookies

This website uses cookies to ensure you get the best experience on our website.