US patent US12748734
Canonical transformations using machine learning language model
Abstract
Various embodiments of the present disclosure provide machine learning techniques for transforming disparate, third-party datasets to canonical representations. The techniques include generating, using a machine learning prediction model, a canonical representation for an input dataset. The machine learning prediction model is previously trained using permutative input embeddings for a training dataset based on canonical data entity features, such that each permutative input embedding corresponds to a different sequence of the canonical data entity features. The permutative input embeddings are leveraged to generate a latent representation for the training dataset. The latent representation is combined with a canonical data map to generate an alignment vector, which is refined to generate an output vector for the input dataset. The machine learning prediction model is trained using a model loss generated based on a comparison of the output vector with a corresponding labeled vector.