Hyperscience AI Transparency Report.pdf
Hyperscience AI Transparency Report
1. Human-readable Content
This is the starting point, representing any document or information that is initially in a format understood by humans.
2. Machine Classification at Target Accuracy
The system automatically categorizes the incoming content based on predefined rules to achieve a desired level of accuracy.
3. Machine Identification at Target Accuracy
The system automatically locates and extracts specific pieces of information or data points from the content, aiming for a set accuracy target.
4. Machine Transcription at Target Accuracy
The system automatically converts handwritten or typed text from images or documents into machine-readable text, striving for a specific accuracy level.
5. Machine Validation Against Business Rules
The system automatically checks the transcribed and identified data against predefined business logic and rules to ensure its validity and consistency.
6. Dynamically Adjustable Threshold (Red Dotted Line)
This indicates that the confidence level required for the machine to automatically process a task can be changed; if confidence is below this threshold, the task is routed to a human.
7. Highly Accurate & Complete Machine-readable Data
This is the final output of the process – data that has been accurately processed and is ready for use by other systems.
8. Submission Complete
This signifies that the processing for a particular set of content or document is finished.
9. QA Tasks
Humans perform quality assurance checks on a sample of the processed data to ensure overall accuracy and identify areas for improvement.
10. Improves Machine Performance
Feedback from human interventions (classification, identification, transcription, exceptions, and QA) is used to train and enhance the machine learning models, making them more accurate over time.
11. Human Intervention
Human intervention can be introduced at any stage of the document processing pipeline, including classification, identification, transcription, and validation as indicated by the boxes above.
AI Model Information
Hyperscience employs a wide array of distinct machine learning and deep learning models ("AI") to perform specific data extraction tasks. These models include both pre-trained and fine-tuned models:
| Model Type | Examples | Description |
|---|---|---|
| Pre-trained by Hyperscience | Rotation Detection, Optical Intelligent Character Recognition(OICR), Text Segmentation, Transcription | General-purpose models trained on public or open-source datasets using proprietary methods. |
| Fine-Tuned | Field Locator, Table Locator, Classification Model, Transcription | Models trained or fine-tuned with data explicitly provided by the user. |
Human-in-the-Loop (HITL)
Hyperscience recognizes the critical role of human oversight in ensuring responsible AI deployment. Our "Human-in-the-Loop" (HITL) approach integrates human input at several key stages:
• Training Data Management
Users can manage training data for trained-from-scratch and fine-tuned models.
• Human-in-the-Loop Oversight
When a model's confidence in its prediction is lower than the set threshold, the system prompts human validation.
• Quality Assurance
Human validation (separate from supervision) is used to measure model accuracy and can be configured to dynamically adjust the workload sent to human reviewers.
Model Performance Measurement
Model outcomes are evaluated based on:
• Automation Rate
The percentage of predictions performed by the machine compared to the total work executed (human + machine).
• Accuracy Rate
The percentage of correct predictions out of the total number of predictions made, as measured by QA tasks.
Data Privacy and Security Compliance
Hyperscience is committed to maintaining the privacy and security of user data through configurable policies and robust data handling practices.
Data Handling & Redaction Pipeline
Customer data undergoes systematic redaction pipelines to remove Personally Identifiable Information (PII) before it is considered for model training. Only data governed by explicit, signed agreements with the original data owners is used in our pipeline. These agreements ensure that all data has been appropriately de-identified and excludes personally identifiable or sensitive information.
Model Training & Data Scope
Our AI models are developed and trained in-house using structured, text-based data. We do not use external or user-generated content containing personally identifiable information (PII), except in cases where a model is explicitly trained for a customer using their own data under a contractual agreement. This controlled approach minimizes exposure to common large language model vulnerabilities, such as those identified in the OWASP Top 10 for LLMs.
Security & AI Risk Management
Hyperscience integrates security best practices throughout the software development lifecycle (Secure SDLC), extending to the development and training of our AI models. We proactively address emerging AI security risks
• AI Bill of Materials (AIBOM)
We are taking strides in an AIBOM solution, providing a cutting-edge view of our AI infrastructure, including model details, dependencies, and training data lineage.
• Regulatory & Contractual Compliance
Hyperscience adheres to GDPR and other relevant data protection regulations, ensuring secure data processing in accordance with contractual obligations and industry best practices.
By implementing these security measures and strict data handling protocols, Hyperscience ensures compliance, privacy, and robust AI security across our platform.
Oversight & Risk Mitigation
Hyperscience mitigates potential AI-related risks through:
• Targeted Model Application
Models are focused on specific data extraction tasks, avoiding broad or unintended consequences.
• Explicit User Training Control
Users must actively opt-in to share data for training.
• Deterministic Outputs
For applicable models, the system returns a consistent output for the same input.
• Robust QA Process
Dedicated quality assurance workflows ensure accountability and safety.
Hyperscience utilizes a variety of datasets for training and evaluating our models, ensuring compliance and responsible data handling:
| Dataset Category | Source | Use Case |
|---|---|---|
| Public Datasets | Research repositories | Data sourced from publicly available research repositories used for benchmarking and evaluation of models. |
| Synthetic & Internal | Generated by Hyperscience | Data generated in-house by Hyperscience for controlled testing, validation, and system performance evaluation. |
| Client & Partner Data | Via explicit agreements | Used solely to train models for the originating customer, in accordance with explicit agreements. |
| Redacted Customer Datasets | Anonymized customer data | Anonymized customer data is used for internal research and model evaluation. |
By providing this detailed overview, Hyperscience aims to foster transparency and trust in our AI-powered platform. We are committed to continuous improvement in our AI practices and will update this report as needed.
Glossary
AIBOM
AI Bill of Materials
A framework that provides detailed information about AI model components, dependencies, and training data lineage.
AI
Artificial Intelligence
Computer systems designed to perform tasks that normally require human intelligence, such as data extraction and decision-making.
HITL
Human-in-the-Loop
An approach where human oversight and intervention are integrated at key stages of AI model processing to ensure accuracy and responsible outcomes.
LLM
Large Language Model
A type of AI model trained on massive datasets of text to understand and generate human language.
OICR
Optical Intelligent Character Recognition
An advanced technology that reads and interprets text from images or scanned documents.
OWASP
Open Web Application Security Project
An organization that publishes guidelines and best practices for web and application security, including risks related to AI and LLMs.
PII
Personally Identifiable Information
Any data that can be used to identify a specific individual, such as names, social security numbers, or contact details.
QA
Quality Assurance
Processes and tasks performed by humans or machines to verify the accuracy and quality of AI outputs.
SDLC
Software Development Life Cycle
A structured process followed for planning, creating, testing, and deploying software.