Trainer

Training Models

What is the Trainer?

The trainer runs various resource-intensive model-training jobs that, if run on the same machine as the application, would slow down submission processing. The trainer runs separately from the main application but connects to it via the API.

Model Management

The Model Management page allows you to see a list of all models trained on this instance. In this article, you'll learn how to navigate the pages for different types of models. To access the Model Management page, go to Library > Models.

Training Data Management

Training Data Management (formerly Keyer Data Management) allows users to improve and supervise models by working directly with the training data (“ground truth”) obtained from each document in the training set. Users can group their documents, select certain documents, and manage data.

TDM for Classification

Classification models are a crucial part of document processing, as they help the system determine which layout should be used to process each page you upload. Training Data Management (TDM) for Classification allows you to add, remove, and update training data.

TDM for Identification Models

Training Data Management (formerly Keyer Data Management) includes tools for controlling and managing Identification models performance. In this article, you will learn how to navigate TDM for Identification Models and its features.

Training Data Curator

Having a diverse, representative training set is crucial for a high-quality identification model. The Hyperscience application allows you to train a model with fewer annotations with minimal impact on performance.

Document Eligibility Filtering

Document Eligibility Filtering indicates whether a document is eligible for training, based on internal checks in the application and our machine learning logic. It provides additional information about documents that were excluded from the training.

Labeling Anomaly Detection

A high-quality model requires consistent annotations. That's why identifying potential discrepancies in the training sets before model training is crucial. To help with this effort, we've included a tool called Labeling Anomaly Detection in Training Data Management.

Text Segmentation

Text Segmentation is the process of partitioning an image into regions containing text into meaningful and distinct pieces or blocks of text. It is the first step of downstream processing tasks such as classification, text transcription, fields, tables, etc.

Requirements for Training a New Model

If you create a new Semi-structured layout version, there will be no models immediately available. For optimal layout performance, train a model on the newest layout version. Recall that Identification models are trained at the layout level.

Training a Semi-structured Model

Hyperscience extracts data from documents and converts them into a machine-readable format. We support Structured, Semi-structured, and Additional documents.

Training a New Field Identification Model

There are two ways to train a Field ID model. To manually train and deploy models, go to the Model Details page, and follow the instructions in this article.

Training a New Table Identification Model

A trained Table ID model enables cell-level predictions and automatic table processing. A Table ID model can be trained to automatically identify both gridded and non-gridded tables.

Retraining Existing Models

Using features for Semi-structured documents. This article mentions features used in the processing of Semi-structured documents.

Training a Classification Model

To achieve better automation rates for document classification, a classification model must be trained for each Semi-structured and Additional layout.

Evaluating Model Training Results

Overview: Once a new Field ID or Table ID model is trained or uploaded, you can evaluate the projected automation based on the Field Identification Target Accuracy setting.

Incremental Training

Adding new data to your training set or making minor changes to its annotations may require several iterations of model re-training. Incremental Training helps you build upon your existing identification model without losing previously acquired information.

Managing Transcription Models

Transcription models are collections of fine-tuning models. Select Transcription models from the drop-down menu at the top of the Models page to view all available fine-tuning models in your instance.

Managing Transcription Models Across Flows

To meet the specific automation needs of your various lines of business, you can configure transcription models at the flow level.

Model Compatibility Logic

Each model is only compatible with one Semi-structured layout, but a model is not necessarily compatible with every version of a layout.

Forward-Compatible Models

In v39, models for flows created in v37 and above are forward compatible, meaning you can use them in v39 without having to retrain them during the upgrade process.

Importing and Exporting Training Data

To ensure that you do not lose any training data during application upgrades and model setups, you can move your training data between environments.

Canceling or Retrying a Training Job

To take action on Field Locator model training, or any training job, navigate to Administration > Trainer.