TDM for Classification Models

TDM for Classification Models

Classification models are a crucial part of document processing as they help the system determine which layout should be used to process each page you upload. Training Data Management for Classification allows you to add, remove, and update training pages for Semi-structured Classification (also known as NLC) models to achieve more accurate classification results. In this way, TDM helps you maximize the performance of your Classification models.

NLC (Non-structured layout classifier) finds the correct Semi-structured or Additional layout for a given set of submission pages, based on the words in the submitted documents. Note that NLC works on a page level. Learn more in Semi-structured Document Classification.

Each release contains a set of layouts. The creation of a release generates a single Classification model. For example, if you create a release with two layouts, then one Classification model will be generated. It will be trained to identify the document pages submitted through the release’s flow.

TDM for Classification logic

TDM for Classification operates on a document level. However, the Training Data tab displays the number of uploaded documents and the required and recommended number of pages per layout. Learn more in the Training Data Tab section of this article.

TDM for Classification allows you to manage example documents that should be included in or excluded from your model’s training:

Limits, requirements, and recommendations:

You need at least 10 page examples to meet the minimum requirements for model training. Our recommendation for a robust model is 120 page examples per layout. Do not upload the same document multiple times.

Access TDM for Classification

To access TDM for Classification, go to the Models section. The Classification tab appears by default.

A table with all Classification models appears:

The Classification models table contains the following columns:

Using TDM for Classification

Using submission data in TDM. The Send documents to Training Data Management setting for Identification and Classification models allows you to control whether submission data is used for model training. It is disabled by default and can be managed from the System Settings (Administration > System Settings).

Overview tab

Projected Automation Chart

The Projected Automation chart displays the performance of the model that’s currently live. Learn more about these metrics in our Accuracy article.

The chart displays how the target accuracy affects the automation. The lower the accuracy, the higher the automation, and vice versa. To learn more, see Automation.

Note that projected model performance (i.e., accuracy and automation) can increase by adding more QA records. You can also see the margin of error (MoE) for this model.

Model History table

The Model History table provides a comprehensive overview of your model’s lifecycle. It displays the following columns:

You can find specific records in the table in the following ways:

Additionally, you can choose which columns are included in the table by clicking the menu next to the Filter drop-down list and clicking the Manage columns… option. You can also adjust the target accuracy by clicking the up and down arrows located next to the Manage Columns option.

Model Compatibility table

The Model Compatibility table indicates the releases that your model is compatible with and contains the columns described below.

You can use the pagination options at the bottom of the table to display all releases that are compatible with your model.

Training Data tab

Training Data Summary card

The Training Data Summary card displays insights on the status of your training dataset.

You need to upload at least 10 examples per layout for the model to learn what documents should be considered as part of your training set. Note that you can run a training without the Excluded examples if you have more than two layouts. However, we recommend adding documents in the Excluded section as well, as they serve as counter-examples.

Training Data Health card

The Training Data Health card displays a breakdown of your dataset. It shows all layouts included in the Classification model, as well as bars next to each layout indicating the number of uploaded pages. Note that you’ll have the required and recommended number of pages for each layout.

Follow the steps below to add training data to your model:

  1. Click the Add Training Data button.
  2. Select the layout you want to add data to from the Upload To Layout drop-down.
  3. Drag and drop your files into the dialog box or click Browse.
  4. Once you’ve uploaded your files, click Continue.

Training Data table

The Training Data table displays all documents that can be used as training data for the model.

It contains the following columns:

Excluded Training Data table

The Excluded Training Data table displays the documents used as counter-examples for your Classification model. The columns are the same as those described above for the Training Data table.

Training a Classification Model

Follow the steps described below to learn how to train a classification model using TDM.

Upload your documents

To upload documents to TDM for Classification:

  1. Click the Add Training Data button on the right-hand side of the Training Data Health card.
  2. Choose Upload Files or Import Training Data from the dialog box.

Review your documents

Review your documents using the Training Document View. It helps you match each document to a specific layout.