# [v42.3] TDM for ORCA VLMs

- Updated on Jul 7, 2026  
- Published on Mar 25, 2026  
- 5 minute(s) read

Training Data Management (TDM) is where you prepare and manage the data used to train your models. In TDM, you review and annotate documents, build your training dataset, and improve model performance for your specific use case.

## Prerequisites

Before you start, ensure the following requirements are met:

- The ORCA base model is installed. Follow the steps in [[v42.3] Installing ORCA VLMs](https://help.hyperscience.ai/v42/docs/v423-installing-orca-vlms) to install and configure the base model in your instance.
- A Semi-structured layout with fields is locked and associated with the flow that is using ORCA. Learn more in [Creating Semi-structured Layouts](https://help.hyperscience.ai/v42/docs/creating-semi-structured-layouts).
- A model definition exists for the layout. The model definition links the layout to the training configuration and enables model training. Learn more in [[v42.3] Model Definitions](https://help.hyperscience.ai/v42/docs/v423-model-definitions).

Learn how to navigate TDM for ORCA VLMs and understand its key sections below.

## Model details page

The Model details page is where you manage and evaluate your model. It includes three main tabs:

- Overview
- Training Data
- History

### Overview

The **Overview** tab provides a summary of the deployed model, the candidate model, the health of the training data, and projected automation.

The sections below explain the key information shown in the Overview tab.

#### Model summary card

The **Model summary** card provides insights into all models trained for a specific model definition. See the table below to learn about the displayed details:

| Field                     | Description                                                                                       | Notes                                                                                                                                                                                           |
|---------------------------|---------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **State**                 | Shows the model’s state.                                                                          | - A **Live** model is a deployed model.<br>  <br>- A **Candidate** model is an undeployed model. It allows you to review model performance and decide whether the model should be deployed or retrained with additional data.<br>  <br>- An **Inactive** model is a candidate or an undeployed model. |
| **Projected automation**   | Displays the performance of the model that’s currently live.                                     | The predicted automation based on the desired target accuracy. The projection is derived from the model’s training data. The system automatically ensures that the same data is not used for both projections and training.  |
| **Test Target Accuracy**   | The accuracy percentage used to calculate the projected automation.                              | Indicates the desired overall system accuracy.                                                                                                                                                  |
| **Trained**               | Date the model was trained.                                                                      |                                                                                                                                                                                               |
| **Layout version**        | The layout version used for this model definition.                                               | Always use the latest locked layout version.                                                                                                                                                  |

#### Training data card

The Training data card displays information about your dataset. See the table below for more information:

| **Field**               | Description                                                                                         | Notes                                                                                                                                                                       |
|-------------------------|-----------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **Training status**     | Indicates the status of your model based on the training data.                                     | - **Ready to train**<br>  <br>- **Reqs not met**<br>  <br>- To meet the requirements for training a specialized model on top of a base model, you must:<br>    <br>- Install the base model. Learn how to do this in [[v42.3] Installing ORCA VLMs](https://help.hyperscience.ai/v42/docs/v423-installing-orca-vlms).<br>      <br>- Annotate at least 120 documents to start training. Learn how to train a specialized model in our [[v42.3] Training a Specialized Model](https://help.hyperscience.ai/v42/docs/v423-training-a-specialized-model) article.<br>- **Training failed** — the model training process failed. <br>  <br>- **Training in progress**— model is currently training. |
| **Total documents**     | The number of annotated documents for this model.                                                   |                                                                                                                                                                             |

### Projected Automation

The Projected Automation chart displays the performance of the currently live model, compared to the candidate one.

The chart displays how the target accuracy affects the automation. The lower the accuracy, the higher the automation, and vice versa.

### Training data

The Training Data tab lists all training documents and allows you to annotate and manage them.

### VLM Annotation

#### VLM annotation experience

While the ORCA VLM works out of the box, you can improve its performance by training it on your specific data, using annotated documents. Learn how to annotate documents.

### History

The History tab displays all models trained for that model definition. This tab allows you to deploy, undeploy, and reject your models. You can also see detailed information for each model.

#### Training results

After training completes:

- a candidate model appears in the Overview tab
- the system calculates projected automation
- the candidate model can be reviewed and deployed.

You can retrain the model by adding more annotated documents and running training again.

#### Deploying the candidate model

To start using the trained model:

- Open the **Overview** tab.
- Review the candidate model summary.
- Click **Deploy** from the **Actions** drop-down to promote the candidate model to Live.

## Next steps

Train a specialized model for your specific use case on top of the ORCA base model and evaluate it, by following the instructions in [[v42.3] Training a Specialized Model](https://help.hyperscience.ai/v42/docs/v423-training-a-specialized-model).

A foundational model that provides core general-purpose capabilities and is not directly trained on customer-specific examples. Use-case specialization is achieved through additional training on top of the base model using customer-specific data.
