Proven Performance: Hyperscience Outperforms LLMs, Open Source, and Legacy IDPs - Hyperscience

Proven Performance: Hyperscience Outperforms LLMs, Open Source, and Legacy IDPs

"The issue with open source benchmarks… is that they’re often skewed toward a very specific set of uses cases, which are often not actually what any normal person does in your product… I think a lot of them are quite easily gameable."

–Mark Zuckerberg on Dwarkesh Patel Podcast

October 1, 2025

What Benchmarks Can Miss

Benchmarks are a valuable way to establish a common standard to compare solutions, track progress over time and drive innovation. However, as the quote above highlights, organizations can end up stacking the deck in their favor by optimizing for the known, tested scenario rather than for a diverse set of real-world settings.

One way organizations skew benchmark results is by using public test sets which can be seen and known ahead of time. This allows developers, either knowingly or accidentally, to overfit their models to perform well on those specific examples.

The Hyperscience Approach to Benchmarking

At Hyperscience, our global machine learning (ML) team brings years of experience in building and testing models for a wide range of enterprise scenarios. A key element that helps our research team continually improve our platform is running intelligent document processing (IDP) experiments and benchmark testing on proprietary datasets that are unfamiliar to LLMs/VLMs. When benchmarking LLMs that are only publicly available, we utilize public datasets. This dual approach delivers more objective results and better simulates real-world performance and scenarios.

Our benchmarks cut through industry buzz and punchy headlines to deliver substance and real-world tested results. The principles below guide how we design and run tests to ensure they generate meaningful insights for our R&D teams while giving customers, partners, and the market up-to-date, trustworthy data on the state of document processing ML models.

Comparative Document Processing Accuracy: Hyperscience vs. Leading Models

The ML benchmarking team recently tested Hyperscience models against the most well-known LLMs, as well as against models from traditional document processing vendors. In both categories, Hyperscience clearly outperformed the alternative models, delivering industry-leading document processing accuracy rates.

Hyperscience vs. LLMs in Document-Specific Use Cases

The table below outlines the accuracy performance percentages of Hyperscience models when compared to some of the most common LLMs.

Bills of Lading Invoices Receipts Government IDs
(h[s]) Specialized GPU (ORCA) 98 94 93 100
(h[s]) Specialized CPU (OICR) 93 93 77 98
Claude 3.7 Sonnet 75 74 82 98
Claude 3.5 Sonnet v2 77 71 49 90
Gemini 2.5 Pro 66 74 86 94
InternVL3-8B 68 67 73 90
NVIDIA Llama Nemotron Nano VL 8B 49 46 80 83
OpenAI GPT4o* * * 76 *

*Hyperscience cannot submit private datasets through OpenAI GPT4o for Bills of Lading, Invoices or Government IDs due to Terms & Conditions.

Hyperscience vs Enterprise AI Platforms: Measuring Accuracy of Printed & Handwritten Text

In the charts below, we’ve removed the LLM startups and compared Hyperscience to models often encountered in enterprise AI extraction. Hyperscalers like Amazon, Microsoft, and Google offer many models. In this benchmark, we compare the ones most commonly used to power their document focused solutions. When compared to both open-source and closed-source models, Hyperscience consistently outperforms, with our CPU and GPU models delivering the highest “exact match” accuracy.

When it comes to documents, there are no models we know of today that are as accurate as Hyperscience for printed and handwritten text. When you factor additional layers of Human in the Loop, orchestration, and accuracy harnessing, Hyperscience’s accuracy percentages only increase compared to other options.

Note: To be considered an exact match, the model must correctly extract the entire word or phrase from the start. For example, if “Cat” is presented, the model only gets credit for returning a response like “Cat” or “cat” both of which preserve the meaning. Other responses like “dogcat”, “CotAT”, or “ca” would be interpreted as incorrect.

What About Multi-Language Accuracy?

Hyperscience also provides multi-language document support for 200+ languages. For any multinational business, the ability to process multilingual documents is of paramount importance. Errors and inaccurate processing lead to costly delays, poor customer experience, and damage to reputation and brand. This is why the world’s leading enterprises turn to Hyperscience for transformational business process automation in whatever language or languages they operate in. Below are the benchmark results when run with Spanish-language models.

What’s Coming Next

Our team runs an ongoing machine learning benchmarking program that is constantly testing new models. Up next, we’ll be looking at the recently released ChatGPT 5.

What do you think of these results? What would you like to see us test next? We’d love to hear from you! Please reach out to learn more.