TDM for Classification Models
TDM for Classification Models
Classification models are a crucial part of document processing as they help the system determine which layout should be used to process each page you upload. Training Data Management for Classification allows you to add, remove, and update training pages for Semi-structured Classification (also known as NLC) models to achieve more accurate classification results. In this way, TDM helps you maximize the performance of your Classification models.
TDM for Classification logic
TDM for Classification operates on a document level. However, the Training Data tab displays the number of uploaded documents and the required and recommended number of pages per layout.
TDM for Classification allows you to manage example documents that should be included in or excluded from your model’s training:
- Layouts eligible to train - These are the layouts that meet the minimum number of pages required for training. To ensure this requirement is met, upload documents that:
- have pages that match your layout and
- are diverse but still represent your layout.
Our recommendation for a robust model is 120 page examples per layout. You need at least 10 page examples to meet the minimum requirements for model training. Do not upload the same document multiple times.
- Excluded documents - TDM uses these as examples of documents that you expect to process but don't want to match. They serve as counter-examples of the documents that your model should not classify.
Access TDM for Classification
To access TDM for Classification, go to Library > Models and click on Classification Models in the drop-down menu at the top of the page.
The Classification models table contains the following columns:
- Model shows the name of your Classification model.
- Compatible Releases indicate the number of releases the Classification model can predict.
- Status shows the model's current state (e.g., Needs Training or Live).
- Training Status indicates the current state of the model training (e.g., Pending, In Progress, Failed, Canceled, or Last trained on [date]).
Using TDM for Classification
Overview tab
Projected Automation Chart
The Projected Automation chart displays the performance of the model that’s currently live.
- Expand it by clicking the arrow button.
The chart displays how the target accuracy affects the automation. The lower the accuracy, the higher the automation, and vice versa. To learn more, see Automation.
Note that projected model performance (i.e., accuracy and automation) can increase by adding more QA records.
Model History table
The Model History table provides a comprehensive overview of your model’s lifecycle. It displays the following columns:
- Name — The name of the last available model for this layout.
- Date Created— Date and time the model was created.
- Version — The specific version of the model that was trained on.
- Source — Indicates where the model was trained—either within the current instance or externally and then uploaded to this instance.
- Docs trained — The total number of documents used for training the model.
- Last Deploy — The last date and hour the model was deployed.
Training Data tab
Training Data Summary card
The Training Data Summary card displays insights on the status of your training dataset.
- Training Data Status indicates your training data's health based on the number of pages uploaded for each layout:
- Requirements Not Met - The minimum number of required pages uploaded for each layout is 10.
- Not Optimized - Hyperscience recommends uploading at least 120 pages to build a robust classification model.
- Ready To Train- This status will be displayed after you’ve reached the minimum required and the recommended number of uploaded pages to start a model training.
Training Data Health card
The Training Data Health card displays a breakdown of your dataset. It shows all layouts included in the Classification model, as well as bars next to each layout indicating the number of uploaded pages. Note that you’ll have the required and recommended number of pages for each layout.
Training Data table
The Training Data table displays all documents that can be used as training data for the model.
- Document ID shows the unique ID number of the document.
- Pages displays the number of pages in the document.
- Layout shows the layout this example corresponds to.
- Usage Rule indicates the way the system will use the specific document for training:
- Always, Auto, Never, Loading, Anomaly.
Excluded Training Data table
The Excluded Training Data table displays the documents used as counter-examples for your Classification model. The columns are the same as those described above for the Training Data table.
Training a Classification Model
Follow the steps described below to learn how to train a classification model using TDM.
Upload your documents
To upload documents to TDM for Classification:
- Click the Add Training Data button on the right-hand side of the Training Data Health card.
Review your documents
Review your documents using the Training Document View. It helps you match each document to a specific layout.
- Assign a layout to the training document from the Layout drop-down menu on the right-hand side of the page.
- Click Save Changes after you’ve classified your document.
Training a Classification Model
Once you’ve reached the requirements and recommendations, you’ll see a message indicating that your model is ready to be trained for the first time in the Overview tab.
- You can run training from either tab by clicking the Run Training button on the upper-right corner of the page.