Training a Semi-structured Model

Training a Semi-structured Model

Hyperscience extracts data from documents and converts them into a machine-readable format. We support Structured, Semi-structured, and Additional documents. To learn how to differentiate between the document types, see Understand Document Types.

Semi-structured use cases

Layouts for Semi-structured documents help identify and extract data from pages that do not have a consistent structure or fixed visual templates. While the information you need to extract (e.g., identification number, address) remains the same, its location may vary and appear under different names or labels. Examples of Semi-structured documents are paystubs and invoices. They contain key pieces of information that are always present, but their placement can vary significantly across different versions of the document.

Use Semi-structured layouts in the following scenarios:

In this article, you’ll learn how to build a robust Semi-structured model corresponding to your business needs by using our Training Data Management tools.

Step 1 - Sampling Documents

Review your documents

Having a diverse, representative training set is crucial for a high-quality Identification model. Selecting the appropriate documents for training will optimize your Semi-structured model’s performance.

Review your fields and columns

Step 2 - Build your layout and add it to a release

Build your layout

Once you’ve determined the information you want to extract, you need to build your Semi-structured layout.

  1. Create your layout by following the steps in Creating Semi-structured Layouts.

Assign to a release

  1. Add your layout to a release by following the steps described in Adding a New Release.
  2. Follow the steps in Assigning a Release to a Flow to match your release to the flow you are using.

Step 3 - Using Training Data Management (TDM)

Use the tools in Training Data Management to control, manage, and adjust the ground truth of your training sets for Identification and Classification models. Before you start:

Upload your documents

  1. Go to the Model Management page for your layout ( Library > Models).
  2. Click Upload Training Documents and upload each document as its own file.
  3. Click Upload in the dialog box.

All uploaded documents will appear on the Training Documents card.

Step 4 - Analyze your data

Training Data Analysis allows you to group your training documents and receive recommendations to improve the quality of your dataset.

Running Training Data Analysis

We recommend running training data analysis once you’ve uploaded your documents. The system will create groups based on the similarity of your training documents which improves the efficiency of the annotation process.

Receive insights for improving your training data by clicking the Analyze Data button, located in the Training Data Health card.

Analysis results

The results show you the eligibility and importance of each document.

Step 5 - Annotate your documents

Consistent annotations are crucial for a high-performance locator model.

Best practices

  1. Once you analyze the data, you’ll be able to annotate by group. Doing so provides you with more control over the dataset.
  2. After annotating 2-3 documents per group, you’ll be able to use guided data labeling.
  3. Follow the general rule for annotating: left to right, top to bottom.
  4. Make sure to maintain consistent annotations for your fields or columns.

Field Identification

  1. Annotate fields with Multiple Occurrences only when multiple instances of a field are present.
  2. Use multiple bounding boxes when a text is logically connected.
  3. If you don’t see a value for a field (i.e., the field is blank), do NOT annotate it.

Table Identification

  1. When annotating a table, make sure to select a row where all data is present.
  2. Always find your table's first and last rows and ensure they are properly annotated.
  3. Always press the ESC button before submitting a table to ensure the annotations are correct.

Next steps

Step 6 - Review your flow’s settings and train your model

Review the flow’s configurations:

For more precise control over the process, you can configure your flow’s settings.

Run Training

Initiate a model training by clicking the Run Training button.

Step 7 - Evaluate the training results

Deploy your model

Once the model training is complete, you’ll find the candidate model in the model details page.

Evaluate the performance Use the documents you’ve chosen for testing purposes to evaluate the performance of your model.