Abstract
Synthetic data are increasingly used in medical imaging and other data-scarce domains, yet existing approaches focus primarily on realism or downstream utility, leaving the problem of dataset validation underexplored. This paper introduces KnowGen, a framework that models synthetic dataset construction as inference over latent acceptability, defined as the fitness-for-use of generated instances under structured specifications.
KnowGen treats synthetic generation as a specification-driven process and maps each instance to structured evidence capturing semantic alignment and redundancy. Acceptability is modeled as a latent variable inferred from observable features and calibrated supervision, enabling uncertainty-aware validation decisions. The framework further incorporates dataset-level quality functionals that capture coverage, diversity, and non-redundancy, extending validation beyond individual samples to global dataset fitness.
By repositioning synthetic data as calibration substrates for estimating acceptability under limited ground truth, KnowGen provides a principled foundation for constructing verifiable synthetic datasets and advances the study of data and information quality as an inferential, auditable process.