Training Data Management

Training Data Management

Training Data Management allows you to improve and supervise models by working directly with the training data (“Ground truth”) obtained from each document in the Training Set. You can group documents, see incompatible ones, annotate representative parts of them, and detect potential inconsistencies.

The performance of your models depends on the quality of the pages, the diversity of the documents, and the consistency of the annotations. For more information on model-training results, see Evaluating Model Training Results.

TDM includes tools for controlling and managing the Identification and Classification models’ performance. Learn more about model performance in Monitoring Model Performance and Improving Model Performance.

TDM for Identification models

TDM for Identification models includes the following features:

It provides the following key capabilities:

Learn how to use these features to maximize the performance of your identification model in our Training an Identification Model article.

TDM for Classification

TDM for Classification models allows you to add, remove, and update training pages for Classification models. Learn more in TDM for Classification.

TDM for VLM Extraction

Training Data Management (TDM) for VLM Extraction models is where you prepare and manage the data used to train a specialized models on top of the ORCA base model. Learn more in TDM for ORCA VLMs.

Accessing Training Data Management tools

If you have the View Training Data permission (given to System Admin and Business Admin permission groups by default), you can access the Training Data Management tools for a model. Learn more in Permission Groups.

  1. Go to the Models section. Learn more about the models table in Models Page.

  2. Click on the specific tab to view the models you need:

  1. Click on the name of the model you would like to view training data for:

Continuous Model Training

When you import an ID or a Classification model from another instance while Continuous Field Locator model improvement and/or Continuous Classification model improvement are enabled, the model’s automation rates may decrease.

Models only learn from the training data available in their current instance. If the new instance contains limited or no training data, the imported model may be replaced by a lower-performing version.

To maintain optimal performance:

Manually annotated data used to train our machine learning models. We use a subset of this data to assess the performance of your models

A dataset used to teach the system how to recognize and extract information. It includes documents with labeled fields so the system can learn from real examples.

A foundational model that provides core general-purpose capabilities and is not directly trained on customer-specific examples. Use-case specialization is achieved through additional training on top of the base model using customer-specific data.