Training Data Management

Training Data Management

Training Data Management allows you to improve and supervise models by working directly with the training data (“Ground truth”) obtained from each document in the Training Set. You can group documents, see incompatible ones, annotate representative parts of them, and detect potential inconsistencies.

The performance of your models depends on the quality of the pages, the diversity of the documents, and the consistency of the annotations. For more information on model-training results, see Evaluating Model Training Results.

TDM includes tools for controlling and managing the Identification and Classification models’ performance. Learn more about model performance in Monitoring Model Performance and Improving Model Performance.

TDM for Identification models

TDM for Identification models includes the following features:

Starting in v43.1, tags are case-insensitive. Tags with the same name but different capitalization are automatically merged into a single tag. This helps prevent duplicate tags and improves consistency when organizing training documents.

Learn how to use these features to maximize the performance of your identification model in our Training an Identification Model article.

TDM for Classification

TDM for Classification models allows you to add, remove, and update training pages for Classification models. Learn more in TDM for Classification.

TDM for VLM Extraction

Training Data Management (TDM) for VLM Extraction models is where you prepare and manage the data used to train specialized models on top of the ORCA base model. Learn more in TDM for ORCA VLMs.

Accessing Training Data Management tools

If you have the View Training Data permission (given to System Admin and Business Admin permission groups by default), you can access the Training Data Management tools for a model. Learn more in Permission Groups.

  1. Go to the Models section. Learn more about the models table in Models Page.

  2. Click on the specific tab to view the models you need:

    • Classification
    • VLM Field Extraction— Learn more in TDM for ORCA VLMs.
    • Identification
    • Text Classification
    • Transcription
  3. Click on the name of the model you would like to view training data for:

    • For ID models:
      • Click the Field Identification or the Table Identification tab, depending on the type of training data you would like to view.
      • The Training Data Management tools are located on the Training Data Health card.
    • For Classification models:
      • Click on the Training Data tab to edit the documents used for training.

Continuous Model Training

When you import an ID or a Classification model from another instance while Continuous Field Locator model improvement and/or Continuous Classification model improvement are enabled, the model’s automation rates may decrease.

Models only learn from the training data available in their current instance. If the new instance contains limited or no training data, the imported model may be replaced by a lower-performing version.

To maintain optimal performance: