No. 1056 - Generating and validating synthetic unbalanced panel data: evidence from Italian firm microdata
This paper examines methods for generating synthetic data, i.e. artificial datasets that reproduce the statistical properties of the original data while reducing the risk of re-identification, namely the possibility of linking data back to the entities they refer to, thereby breaching confidentiality. We apply the two most widely used approaches in the literature, sequential modelling and Gaussian copulas, to microdata from Banca d'Italia's Survey of Industrial and Service Firms.
Both approaches accurately reproduce the multivariate structure and temporal dependencies of the original data. Greater similarity between synthetic and original data is associated with a higher risk of re-identification, which can be mitigated by introducing random perturbations into the synthetic data, with only a limited loss of analytical usefulness. These techniques therefore appear promising for reconciling the sharing of microdata for research purposes with the protection of statistical confidentiality.
Full text
-
17 September 2026
Instagram
YouTube
X - Banca d'Italia
Linkedin