US patent US12724832
Automatic field discovery for training data extractors
Abstract
Systems and methods are disclosed for automatic discovery, labeling, and extraction of data fields related to submissions such as applications, inputs for calculations, audits, or similar forms. The system extracts values for data fields to populate a data model. An automatic field discovery system receives submission documents and a bill of data specifying data fields to extract and constraints for each field. A discovery agent pool, using multiple language models, searches portions of the documents to generate candidate values with references. A ranking agent pool evaluates the candidates against the bill of data and produces ranked lists with explanations. A synthesis agent reconciles the rankings to select final values, which are stored as training values and combined with submissions to form training samples. An extractor generation system uses the training samples to generate or refine lightweight extractors that a data extraction manager applies to new submissions.