Training Data Curator

Training Data Curator

Having a diverse and representative Training Set is essential for building a high-quality identification model. The Hyperscience Platform helps you identify which documents are most valuable for training, allowing you to prioritize annotations more efficiently.

How data is curated

The Training Data Curator labels each training document as having high or low importance. The importance is calculated by determining which data would best contribute to the model’s performance. For each group of documents, the system labels the most impactful ones as having high importance.

Before using the Training Data Curator, make sure that you've:

Grouping Logic

Documents are grouped based on text and location. Note that new groups will appear every time you run Training Data Analysis. To learn more, see Step 4 of our Training a Semi-structured Model article.

Using the Training Data Curator

  1. Analyze the data.

The Training Data Curator depends on the results from the document grouping. That's why you should first analyze your data.

For more information about data analysis and groups, see Labeling Anomaly Detection.

After data analysis is finished, the Training Data Health card shows the quality of your training set.

  1. Filter the documents by group
  1. Annotate the documents with High importance first.

The more representative your training documents are of real-life documents, the better your model will perform with fewer annotations.

All annotated documents will be marked as high importance during training data analysis, while similar ones will be marked as low importance. These updates help keyers annotate documents that are most valuable to the model. If you import documents that have already been annotated, they have high importance by default.

Next steps

A dataset used to teach the system how to recognize and extract information. It includes documents with labeled fields so the system can learn from real examples.

The suggested minimum and optimal amount of labeled data required to train machine learning models in Hyperscience.

Providing diverse, high-quality examples ensures better model accuracy and reliability.

A tool in TDM that analyzes your training data to compute the importance of each training document and identify issues such as missing labels, overlapping fields or columns, or inconsistent annotations. This analysis helps you prioritize which documents to annotate and ensures clean, accurate data before you train a model.