Training Data Curator

Training Data Curator

Having a diverse, representative training set is crucial for a high-quality identification model. The Hyperscience application allows you to train a model with fewer annotations with minimal impact on performance.

How data is curated

The Training Data Curator labels each training document as having high or low importance. The importance is calculated by determining which data would best contribute to the model’s performance. For each group of documents, the system labels the most impactful documents as having high importance.

Before using the Training Data Curator, make sure that you've:

Using the Training Data Curator

  1. Analyze the data.

The Training Data Curator depends on the results from the document grouping. That's why you should first analyze your data.

For more information about data analysis and groups, see Labeling Anomaly Detection.

After data analysis is finished, the Training Data Health card shows your training set's quality.

The importance of each document appears in the Training Data Table. Importance is partially determined by grouping or document similarity—similar documents are grouped together and assigned high or low importance, helping you avoid redundant annotations.

  1. Click Filters to filter the documents by group and by importance.

After you’ve selected a group and importance, Apply Filters.

.jpg?sv=2026-02-06&spr=https&st=2026-07-27T09%3A02%3A23Z&se=2026-07-27T09%3A15%3A23Z&sr=c&sp=r&sig=BhMv2DDf6pCIYlUtI%2Bgk0wJe0RaL5p3B7MMz%2FF8j2Ko%3D)

  1. Annotate the documents with High importance first.

The more representative your training documents are of real-life documents, the better your model will perform with fewer annotations. We recommend double-checking the ones with Low importance.

All annotated documents will be marked as high importance during training data analysis, while similar ones will be marked as low importance. These updates help keyers annotate documents that are most valuable to the model. If you import documents that have already been annotated, they have high importance by default.

Next steps