Detecting and Correcting Anomalies in Table Annotations

Detecting and Correcting Anomalies in Table Annotations

A high-quality model requires consistent annotations. That's why the Hyperscience application has a tool, called Labeling Anomaly Detection, for identifying potential discrepancies in the training datasets before running model training. Once the annotations are ready, the user can analyze the data to find inconsistencies and ensure a top-performance locator model.

In v37+, this tool is available for both Field Identification and Table Identification models. Learn more about Field anomaly detection in Detecting and Correcting Anomalies in Field Annotations.

Limitations of Table ID Labeling Anomaly Detection in v37 and above

Detecting anomalies

Labeling Anomaly Detection now includes anomalies generated from Model Validation Tasks after training in previous versions. For more information, see Model Validation Tasks.

Before using Labeling Anomaly Detection, make sure that you've uploaded the required number of training documents.

  1. Go to Library > Models, and make sure Identification Models are selected from the drop-down list at the top of the page.

  2. Find the model you want to work on, and click on its name to access its Model Details page.

  3. On the Model Details page, click on the Table Identification tab.

  4. In the Training Data Health card, click Analyze Data.

  5. Re-analyze your data after annotating the documents with High importance.

If anomalies were detected during the analysis:

Reviewing Anomalies

  1. Above the Training Data table, click Filters, and then select a group from the Group drop-down list.
  2. Click Apply Filters.
  3. Click the Edit Annotations link for a document highlighted as having potential anomalies.

The right-hand sidebar shows the potential anomalies.

  1. Click on a column to access its action buttons.

You can select columns by clicking on the colored labels or the column names on the right-hand sidebar. You can also use the W and E keys on your keyboard.

  1. Do one of the following:
    • Make the required correction (adjusting the bounding box, annotating the missing cell, etc.) When you do so, the anomaly disappears.
      • Annotate any missing columns to remove the “Missing column” label.
      • If the column is not present in the document, click on the checkmark in the column’s colored marker to remove the label.
      • If a nested table contains anomalies, click on the name of the parent or child table in the right-hand sidebar. Then complete steps 4-5.
      • Learn more about nested tables in Table Identification.
    • If the anomaly appears correct, click on the checkmark in the column’s colored marker to remove it.

Anomaly indicators

The indicators described in the table below appear in the document viewer if anomalies are detected in the document.

Indicator Description Example
Cell-anomaly indicator

Always make sure to reanalyze your data for updated information on the training set. If documents have been added, removed, or modified since the last analysis, the ineligibility details may be outdated.