TDM for Classification

TDM for Classification

Classification models are a crucial part of document processing, as they help the system determine which layout should be used to process each page you upload. Training Data Management (TDM) for Classification allows you to add, remove, and update training pages for Classification (also known as NLC) models to achieve more accurate classification results. In this way, TDM helps you maximize the performance of your Classification models.

NLC (Non-structured layout classifier) finds the correct Semi-structured or Additional layout for a given set of submission pages, based on the words in the submitted documents. Note that NLC works on a page level. Learn more in Automatic Document Classification.

Each release contains a set of layouts. The creation of a release generates a single Classification model. For example, if you create a release with two layouts, then one Classification model will be generated. It will be trained to identify the document pages submitted through the release’s flow.

TDM for Classification logic

TDM for Classification operates on a document level. However, the Training Data tab displays the number of uploaded documents and the required and recommended number of pages per layout. Learn more in the Training Data Tab section of this article.

TDM for Classification allows you to manage example documents that should be included in or excluded from your model’s training:

Our recommendation for a robust model is 120 page examples per layout.

You need at least 10 page examples to meet the minimum requirements for model training.

Do not upload the same document multiple times.

Using TDM for Classification

Learn how to navigate through TDM for Classification and how to use it in the tabs below.

Overview

Upload and Review

Training a Classification model

Classification Models Table

Access TDM for Classification

To access TDM for Classification, go to Library > Models, and click on Classification Models in the drop-down menu at the top of the page.

A table with all Classification models appears:

Importing Classification Models

When importing classification models, make sure they are trained for the same release version as the currently opened classification model.

If the model is from a different release, an error message will appear in the UI indicating the mismatch.

The Classification models table contains the following columns:

To access TDM features for your Classification model:

  1. Go to Library > Releases
  2. Click on the release's name for the model you want to manage training data for.
  3. Click View Model in the Automatic Document Classification card.

Click on the Overview or Training Data tab to learn more about the model and optimize its performance.

Overview tab

Classification Overview tab

The Classification Overview tab contains the tables described below.

Projected Automation Chart

The projected Automation chart displays the predicted automation rate of your model based on the target accuracy. Learn more about these metrics in our Accuracy article.

Model Activity table

The Model Activity table shows the training history for your classification models and has the following columns:

You can also download the current version of your model from the drop-down menu next to the Run Training button.

If you download the training data for the model, it may contain personally identifiable information. Learn more about managing your data in PII Data Deletion.

The System Version is the Hyperscience version the model was trained in. Change it from the System Version drop-down menu.

Use the pagination options at the bottom of the table to display all activities for your Classification model.

Model Compatibility table

The Model Compatibility table indicates the releases that your model is compatible with and contains the columns described below.

You can use the pagination options at the bottom of the table to display all releases that are compatible with your model.

Training Data tab

Training Data tab

The Training Data tab contains the sections described below.

Summary

This section shows your training data stats, as well as the date and hour of the last model training.

You need to upload at least 10 examples per layout for the model to learn what documents should be considered as a part of your training set. Note that you can run a training without Excluded examples if you have more than two layouts. However, we recommend adding documents in the Excluded section, as well, as they serve as counter-examples.

Training Data Health

The Training Data Health card displays a breakdown of your dataset. It shows all layouts included in the Classification model, as well as bars next to each layout indicating the number of uploaded pages (not documents, as suggested in the application). Note that you’ll have the required and recommended numbers of documents for each layout.

Training Data

The Training Data table shows all documents available for use as training data for the model. It contains the following columns:

Hover over the preview icon to see the pages of the uploaded document. Freeze the preview by clicking on the preview icon. Page through the document using the arrow keys on your keyboard or the arrows in the preview dialog. Click anywhere on the page to hide the preview.

Note that TDM for Classification works on a document level (i.e. when you edit the classification in TDM, you will classify the whole document and not a single page to a specific layout), whereas QA operates at the page level. For example, if you classify 3 pages into 2 different layouts in QA, those 3 pages will be combined into a single document in TDM. That document will keep the machine's prediction for the layout.

The machine might classify a page with high confidence yet still be incorrect. This type of mistake is known as a high-confidence error. To confirm and correct such errors, users must complete Model Validation Tasks (MVTs), which are shown as anomalies in TDM. Learn more in Document Classification Model Validation Tasks.

Excluded Training Data

The Excluded Training Data table displays the documents used as counter-examples for your classification model. The columns are the same as those described above for the Training Data table.

You can change the displayed columns by clicking on Manage Columns… in the drop-down menu.

You can filter the tables by:

You can also search by Document ID. The Actions drop-down menu provides options to bulk-delete, edit, or download training data, as well as to download the entire training dataset.

Uploading documents

To upload documents to TDM for Classification:

  1. Click the Add Training Data button on the right-hand side of the Training Data Health card.
  2. Choose Upload Files or Import Training Data from the dialog box.

You can import the following training data:

Training Document View

The training document view helps you match each document to a specific layout.

Note that TDM for Classification works on document-level (i.e. you will classify the whole document and not a single page to a specific layout).

Classification Model Training

Different Classification models across flows share the same data in TDM. Any changes applied to the training data (e.g., updating or removing documents) are also applied to all releases and flows.

After you’ve reached the requirements and recommendations, you’ll see a message that indicates that your model is ready to be trained for the first time in the Overview tab.

You can run training from either tab by clicking on the Run Training button on the upper-right corner of the page.

Classification models are automatically deployed after training.

You can cancel your training at any time from the drop-down menu in the upper-right corner of the page.

Learn more about Classification in Document Classification.