Abstract
Introduction: Experimental tensile testing of sub-sized nuclear structural materials is often limited by the high cost, time requirements, and scarcity of available data. These constraints lead to sparse and imbalanced datasets that challenge the development of reliable predictive models. This study investigates the effectiveness of machine learning (ML) algorithms and data augmentation strategies for predicting tensile properties of sub-sized Stainless Steel 316 specimens.Methods: A dataset comprising 396 tensile testing records and 45 input features describing material composition, processing history, irradiation conditions, specimen geometry, and testing parameters was assembled from the literature. Nine ML algorithms were evaluated for predicting ultimate tensile strength, yield strength, total elongation, and uniform elongation. Two data augmentation methods, Generative Adversarial Networks (GAN) and Synthetic Minority Oversampling Technique for Regression with Gaussian Noise (SMOGN), were applied at multiple augmentation ratios. Model performance was assessed using five-fold cross-validation with Pearson correlation coefficient (r), coefficient of determination (R2), and root-mean-square error (RMSE).Results: Tree-based ensemble methods achieved the highest predictive performance across most tensile properties. Random Forest and Extreme Gradient Boosting (XGBoost) attained Pearson correlation coefficients exceeding 0.98 for yield strength and uniform elongation. Data augmentation effects were property dependent: GAN augmentation generally improved predictions of strength-related properties, while SMOGN yielded greater benefits for elongation-related properties. The best augmentation configurations improved predictive accuracy and model robustness, particularly for sparse regions of the feature space. Physics-Informed Neural Networks incorporating empirical constraints underperformed conventional ML approaches, exhibiting approximately 10% lower correlation and higher prediction errors.Discussion: The results demonstrate that combining suitable ML algorithms with property-specific data augmentation strategies enables accurate prediction of tensile properties from limited experimental data. The differing effectiveness of GAN and SMOGN suggests that augmentation strategies should be selected according to the physical characteristics of the target property. The proposed framework provides a practical approach for addressing data scarcity in nuclear materials research and supports accelerated qualification and evaluation of nuclear structural materials.