Training Data Analysis

Training Data Analysis

Training Data Analysis is a tool in Training Data Management (TDM) that helps you understand the quality of your dataset before training a model. It analyzes the uploaded documents to identify patterns based on text and location. The details you can see from the analysis are:

Use Training Data Analysis to:

Running Training Data Analysis before training or retraining helps improve data quality and reduce the risk of poor or unstable model performance. This article explains how documents are grouped and how to use those groups during the annotation process.

Running Training Data Analysis

Training Data Analysis runs on the documents in your training dataset. To begin the analysis:

  1. Upload your dataset to TDM.
  2. Click the Analyse Data button.

.jpg?sv=2026-02-06&spr=https&st=2026-07-27T09%3A14%3A55Z&se=2026-07-27T09%3A27%3A55Z&sr=c&sp=r&sig=85ZrhnEM3lT9VldozTarRGLxH%2BS%2FBfk37Xkapfrhyxk%3D)

Reanalyze data

The analysis evaluates the current state of the dataset and generates groups, importance scores, anomaly indicators, and eligibility results. Because the analysis is relative to the current dataset, results may change each time you rerun it. Make sure to reanalyze the data each time you:

Grouping logic

As described above, Training Data Analysis groups documents based on similarities in text and location. Each group represents a distinct pattern within the current training dataset and helps you determine how documents are distributed. Grouping is relative to the dataset at the time the analysis is run. When documents are added, removed, or updated, and the analysis is rerun, group composition changes. That’s why new groups may appear, existing groups may merge, or documents may shift between groups.

Using groups during annotation

Annotating documents by group improves efficiency in the process and consistency in the training data. Since documents within a group share similarities, annotating them together helps maintain consistent field labeling.

When preparing training data:

Groups and dataset diversity

The number and size of groups provide insights into dataset diversity.

Understanding group distribution helps you assess whether the dataset reflects expected real-world examples. To learn how to investigate model performance issues using the groups, see Improving Model Performance.

A tool used to annotate, manage, import, and export training documents. It is also used to train models by working directly with the training data (“ground truth”) obtained from each document in the training set.

A third-party entity that does business with your company and sends you documents. The meaning of the data can vary depending on the vendor.