Labeling Anomaly Detection

Labeling Anomaly Detection

Accessing this feature
Your access to the feature described in this article depends on your license package and pricing plan.
To learn which features are available to your organization and how to add more, contact your Hyperscience representative.

A high-quality model requires consistent annotations. That's why identifying potential discrepancies in the training sets before model training is crucial. To help with this effort, we've included a tool called Labeling Anomaly Detection in Training Data Management (TDM).

After completing the annotations, you can analyze the Training Data to find inconsistencies in your documents and the ones ineligible for training. To learn more about eligibility, see our Document Eligibility Filtering article. Labeling Anomaly Detection identifies and highlights potential anomalies in field and table annotations for review.

Before using Labeling Anomaly Detection:

  1. Upload the Required Number of Documents (100 minimum, 400 recommended).
  2. Run Training Data Analysis.
  3. Annotate your Training Set.
  4. Reanalyze your data.

Always re-run data analysis to get the most up-to-date information about your training set. Ineligibility details may change if documents have been added, removed, or modified since the last analysis. For more information, see Step 4 of our Training a Semi-structured Model article.

Detecting anomalies

If anomalies were detected during the training data analysis:

  1. Expand the Filter section, and select Contains Anomalies from the Has Anomalies drop-down list.
  2. Click Apply Filters.
  3. Click the Edit Annotations link for a document highlighted as having anomalies.
  4. Review the annotations highlighted as being a potential anomaly.

Limitations of Labeling Anomaly Detection

Anomaly indicators

The indicators described in the table below appear in the document viewer if anomalies are detected in the document.

Indicator Description Example
Cell-anomaly indicator - found in the right-hand sidebar Dotted line around a cell - indicates a cell that needs to be reviewed
Page Indicator - found in the left-side sidebar Dotted line around a page - indicates cell-level anomalies on the specific page
Missing columns label - found on the top of the document, next to the colored markers for each column The indicator for a missing column is a dotted, transparent label, located next to the bookmark indicators for each column.
Missing column tag - found in the right-hand sidebar next to the specific column that is missing This tag shows the specific missing column.
Misplaced column - found around the colored markers for columns Dotted line around the colored markers - indicates misplaced columns
Misplaced column tag - found in the right-hand sidebar next to the specific column that is misplaced This tag shows the specific misplaced column.
Number of anomalies - found in the right-hand sidebar above the list of columns in the document This yellow indicator displays the current number of potential anomalies in the document. It is dynamic and changes after each interaction with an annotation labeled as an anomaly.
If you have a nested table, the number of potential anomalies will also appear next to the name of the parent or child table.
Ignore anomaly - action button, located in the colored label for a column The bell button appears for single and multiple anomalies in a column. Click on it to ignore the detected anomaly
Single anomaly - hover over it to see the specific anomaly
Multiple anomalies - hover over to see the number of potential anomalies for this column

Re-analyzing data

We recommend re-analyzing the data after reviewing all anomalies to ensure the training set is consistent and ready for model training. Click Reanalyze data to choose one of the two options listed below:

A tool used to annotate, manage, import, and export training documents. It is also used to train models by working directly with the training data ("ground truth") obtained from each document in the training set.

The input used to teach machine learning models how to process documents accurately. Its structure depends on the model type:

The minimum amount of annotated data needed to train machine learning models in Hyperscience.

Meeting these thresholds ensures the models can be successfully trained and begin learning layout- or document-type patterns.

A tool in TDM that analyzes your training data to compute the importance of each training document and identify issues such as missing labels, overlapping fields or columns, or inconsistent annotations. This analysis helps you prioritize which documents to annotate and ensures clean, accurate data before you train a model.

A dataset used to teach the system how to recognize and extract information. It includes documents with labeled fields so the system can learn from real examples.

Annotation refers to a user-provided input that defines the correct prediction for a given machine learning task. Annotations are used to train supervised machine learning models.

A rectangular subregion of a given page that specifies the location of text to be processed downstream or to be displayed to the user.

Multiple Occurrences (MOs) are used to identify multiple different instances of a field.

Used to annotate a single field's value when it spans across line or page breaks, ensuring the accurate capture of information across multiple locations or pages.