Labeling Anomaly Detection

Labeling Anomaly Detection

Accessing this feature
Your access to the feature described in this article depends on your license package and pricing plan.
To learn which features are available to your organization and how to add more, contact your Hyperscience representative.

A high-quality model requires consistent annotations. That's why identifying potential discrepancies in the training sets before model training is crucial. To help with this effort, we've included a tool called Labeling Anomaly Detection in Training Data Management (TDM).

After completing the annotations, you can analyze the training data to find inconsistencies in your documents and any documents that are ineligible for training. To learn more about eligibility, see our Document Eligibility Filtering article. Labeling Anomaly Detection identifies and highlights potential anomalies in field and table annotations for review.

Before using Labeling Anomaly Detection:

  1. Upload the required number of documents (100 minimum, 400 recommended).
  2. Run training data analysis.
  3. Annotate your training set.
  4. Reanalyze your data.

Always re-analyze your data for updated information on the training set. The ineligibility details may need to be updated if documents have been added, removed, or modified since or during the last analysis. You can learn more about how to use Training Data Analysis in Step 4 of Training a Semi-structured Model.

Detecting anomalies

If anomalies were detected during the training data analysis:

  1. Above the Training Documents card, click Filters, and select Contains Anomalies from the Has Anomalies drop-down list.
  2. Click Apply Filters.
  3. Click the Edit Annotations link for a document highlighted as having anomalies.
  4. Review one of the annotations highlighted as being a potential anomaly.
    • Check how the field was annotated in other documents in the same group. That way, you'll ensure consistency throughout the training set. If the annotation is not correct, adjust it accordingly and click Save Changes.
    • If the annotation is correct, click on it and then click Ignore Anomaly. A warning message will appear.
    • Click Confirm and then Save Changes.

Limitations of Labeling Anomaly Detection

Anomaly indicators

The indicators described in the table below appear in the document viewer if anomalies are detected in the document.

Indicator Description Example
Cell-anomaly indicator

Re-analyzing data

We recommend re-analyzing the data after reviewing all anomalies to ensure the training set is consistent and ready for model training. Click Reanalyze data to choose one of the two options listed below: