Table Identification

Table Identification

A table is a data structure used to organize and present information in rows and columns. It is used to present values in a readable format.

Each table consists of the following elements:

Hyperscience provides a solution for extracting data from tables by using Table Identification models in Semi-Structured layouts.
We support extraction for the following tables:

In this article, you will learn how to annotate and train a Table ID model.

Table Identification Task

Table ID tasks are specific to Semi-structured documents with table columns. To see these tasks, you must define table columns on a Semi-structured Layout.

Prerequisites

Follow the steps below to define a table and start the annotation process:

Annotating Tables

The Table Identification task is available in Training Data Management, and it’s similar to annotating in Supervision.

Navigating the Table ID task in TDM

Be sure to click the Continue to Review (CMD+ENTER) button for each available table.

Follow the steps below to annotate your table in Supervision:

1. Select a row from the table to be a Template row

A Template Row is the lead row in your table. It is not necessary to be the first one. Hyperscience uses the Copycat tool to populate the annotation to the rest of the rows in Step II. The copycat is not always accurate, so make sure to double-check the annotations.

In this step, you will define the template row of your table.

  1. Select a row from the table to be a template row.

2. Make sure to capture the cell and follow the tips below if necessary

| If... | Then... | | The Bounding Box includes all of the cell's content. | Move on to the next step. | | The box is in the right place, but doesn't include all of the cell's content (e.g., parts of letters fall outside of the box). | Click and drag the box's corners until it contains all of the content that should be transcribed. | | Neighboring text segments should also be included in the cell's transcription. | With a click-and-drag motion, draw a bounding box that includes all of the cell's content. | | The box doesn’t include any of the cell’s content, OR no bounding box appears around the cell’s content when hovering over it. | Press the spacebar, and with a click-and-drag motion, draw a bounding box that includes all of the cell's content. |

3. Review your annotations

If you have more rows in the table, use the Split button to identify them faster:

Other actions

You can use the Scroll freeze button located at the top of the page if you have more pages in your document. Clicking it improves the performance of the system by rendering the images on each page faster.

Quality Assurance (QA) is a process that ensures the accuracy and reliability of system outputs. In Hyperscience, QA tasks allow users to review and correct errors in classification, identification, VLM extraction and transcription.

After you’ve annotated a single row from a table, you can use the copycat feature to copy the annotations to the remaining rows of the table. The copycat is not always accurate, so make sure to double-check the annotations before you submit.

A rectangular subregion of a given page that specifies the location of text to be processed downstream or to be displayed to the user.

The model for each table is available in the Table Identification Models card. To initiate a model training follow the steps described in Training a New Table Identification Model.
Learn more about navigating Training Data Management and using its features in Training Data Management Features and Training Data Management.

A configuration within Hyperscience designed to process documents where fields and table cells are present, but their positions can vary among documents.