Structured Document Classification

Structured Document Classification

Classification models in Hyperscience

Classification models are a crucial part of document processing as they help the system determine which layout should be used to process each page you upload. In Hyperscience, we have two types of document classification:

In this article, you’ll learn how to work with Structured Document Classification. Learn more about Identification in TDM for Identification Models. To learn about Transcription and Flexible Extraction, see our Transcription and Flexible Extraction articles.

How it works

When a document is submitted to the system, it is first split into individual pages. Structured Document Classification then runs through the following steps:

  1. Visual Page Classifier (VPC):
    The system runs VPC, which returns layout page-level candidates for each submission page. At this stage, only candidates are proposed, not full documents.

    • For every page, the VPC produces a list of possible layout matches, ranked by confidence. To learn more about Structured layouts, see Creating Structured Layouts.
    • Example: Submission page 1 may match layout page 1 of Layout A; submission page 2 may match layout page 2 of Layout A, etc.
  2. Distribute to Forms:
    Using the list of candidates generated from the VPC, the system attempts to group pages into complete documents.

    • The goal is to minimize the number of documents while maximizing confidence scores across all pages.
    • Pages that fail to align with any layout at this stage are passed to Semi-structured classification (Non-Structured Layout Classifier (NLC)). To learn more, see Semi-structured Document Classification.
  3. Registration:
    For each proposed distribution, the system runs a Registration step to validate the page-to-layout matches.

    • If confidence for all candidates is above the acceptance threshold (e.g., >0.6), the document is accepted.
    • If one or more candidates are rejected, the system re-runs Distribute to Forms with alternative matches.
  4. Re-distribution & Finalization:
    The Distribution - Registration cycle runs up to three times:

    • First attempt: Initial grouping of candidates.
    • Second attempt: Re-distribution if a candidate is rejected.
    • Third attempt: Final re-distribution.
    • If all candidates are successfully registered, the current distribution is finalized and used.
    • If some candidates are rejected, the system generates a new distribution and retries registration.
    • If none of the attempts result in a fully valid distribution, the system returns the best-scoring distribution from the three tries.
    • This process repeats up to 3 times.
  5. Manual Review:
    Any pages that fail to classify after these steps are marked as No Layout Found and routed to Semi-structured Classification. Depending on your flows settings and your use case, these can be handled via Document Classification Supervision Task. To learn more, see Semi-structured Document Classification.

Structured Layout Match Threshold

The Structured Layout Match Threshold defines the minimum confidence score required for a page to be automatically matched to a Structured layout.

Layout Matching Confidence

Confidence in layout matching directly affects the accuracy. The more confident the system is in its layout match, the more reliable the extracted data will be. To learn more, see our Accuracy article.

Expand the sections below to learn more about the Structured document classification settings and layout identifiers.

Structured Document Classification Settings

Before you start, configure Structured document classification behavior in your flow.

Learn more about these settings in the Classification section of our Document Processing Subflow Settings article.

Layout Identifiers

Classifying Variations

Some layouts can look almost identical, with only minor visual differences. To avoid misclassification in these cases, you can create layout variations. Each variation represents a small difference in the layout’s pattern, while still belonging to the same overall layout group. Learn more in Adding a Variation to a Layout.

Layout Identifiers

Even with variations, the system may sometimes classify incorrectly. To improve accuracy, Hyperscience can use Layout identifiers to force the correct match.

Using Layout Identifiers

  1. Go to Library > Layouts.
  2. Find the layout to which you want to add a layout ID, and click on its name.
  3. Find the variation to which you want to add a Layout ID, and click on its name.
  4. Click Fields in the toolbar, and then click Layout IDs.
  5. Click and drag to draw bounding boxes around each layout ID.
    • Once you draw a box, the machine will read and transcribe the value inside it.
    • Any incorrect transcriptions can be edited in the field list.
  6. When you’re finished making changes to the variation, do one of the following:
    • If you’re ready to apply your changes to the variation, click Commit Changes and save it as a new version. Learn more in Editing and Finalizing a Layout Version.
    • If you’re not ready to apply your changes, click the X button in the upper-right corner of the page.

Layout Identifiers Best Practices

The following best practices for using layout identifiers will ensure the highest levels of accuracy for the classification of Structured documents.

Document Classification task

Document Classification is a Supervision task used to group pages into documents and assign the correct layout. These tasks are created when automatic classification isn’t possible (e.g., when no model is available or the model’s confidence is low).

Document Classification Interface

To open a Document Classification task, go to the Tasks section and click Perform Tasks under the Supervision task type table.

Document Classification allows you to:

Left Panel - Uncategorized

The left panel displays all uncategorized pages - pages that haven’t yet been grouped into a document.

Middle Panel - Grouped Documents

The middle panel shows all documents you’ve grouped. In this panel, you can:

Right Panel - Document Preview

The right panel shows a zoomable full-page view of the selected page.

Reprocessing Misclassified documents

If a document is marked as Layout Incorrect during Flexible Extraction, Identification, or Transcription Supervision, you can reprocess it and send it back to Document Classification for rework.

Using Reprocessing

You can trigger reprocessing from any Identification or Transcription Supervision or Flexible Extraction task. To learn more, see Supervision and QA Introduction.

  1. Go to Submissions.
  2. Find the mismatched submission by its ID.
  3. Click on the Perform Tasks link in the Tasks column for the submission.
  4. In the right sidebar of the task, expand the document information and click Mark Layout Variation Incorrect.
  5. Click Continue on the warning message.
  6. Reclassify your documents by following the guidance in Document Classification Interface section.