---
Source: https://docs.microblink.com/verify/evaluation
Title: Evaluate Verify
Description: Evaluate and test BlinkID Verify performance before production deployment
---

# Evaluate Verify

Before going to production, you may want to evaluate Verify on your own data.
The difficulty of this evaluation depends on the type of data that you have.

## What to expect

On a global evaluation dataset of hundreds of thousands of real-world images of real and fake documents, Verify achieves highly reliable results with the default configuration.

When controlling for image quality, the false rejection rate ([FRR](./glossary.md#frr)) is low both globally and in the USA.

The evaluation also includes liveness false rejections, but those datasets don't have screen/photocopy presentation attacks.
On dedicated liveness datasets, screen/photocopy (and aggregate liveness) false acceptance rate and false rejection rate remain low.[^1]

## False acceptance and false rejection

If you have a sample set of real document images, you can measure Verify's false rejection rate.

But if you have only images of real documents, you can't measure the false rejection rate.
Inversely, if you only have images of fake documents, you can't measure the false acceptance rate.

The balance between FRR and FAR is important because any system can achieve a perfectly low false rejection score by accepting everything (including fraudulent documents), or a perfectly low false acceptance score by rejecting everything (including real documents).

:::tip[Expired documents]
Be careful when testing with expired documents!
By default, Verify rejects expired documents.
If your dataset contains expired documents you want to treat as genuine, set [`rejectExpiredDocuments`](./configuration.md#expiration) to `false`.
:::

In theory, you can simply disable all checks and get 0% FRR.
Without fake documents, you won't know you're not catching any fraud.
Similarly, you might be comparing two products, and if one has 1.5% FRR, while the other one has 1.75% FRR, you might conclude the first one is better.
But what if it catches 10 times less fraud then the second one?

We recommend acquiring a quality dataset of synthetic fake documents, like the Department of Homeland Security's [IDNet dataset](https://www.researchgate.net/publication/382884673_IDNet_A_Novel_Dataset_for_Identity_Document_Analysis_and_Fraud_Detection).

## Recommended sample size for evaluation

To ensure that the evaluation results accurately reflect BlinkID Verify's performance, it is important to use a large enough and diverse sample set.
Evaluations based on a very small number of images (for example, 5 examples per document type) can lead to misleading conclusions due to lack of representative data.
We recommend the following sample sizes for reliable testing:

Minimum setup:

- 20 to 30 real document examples per document type
- 20 to 30 fake document examples per document type

Recommended setup:

- around 100 real and 100 fake examples per document type

This way, your evaluation results will more accurately reflect real-world performance.

## Trade-offs

FRR and FAR exist in tension.

If you want to increase the number of accepted real documents, the trade-off is that you will also increase the number of accepted fake documents.

If you want to increase the number of rejected fake documents, the trade-off is that you will also increase the number of rejected real documents.

It's important to identify where you want to be on this scale, and [tune the solution](./configuration.md) to get there.
This is typically done by fixing one of the two metrics.

For example, you might know you don't want to reject more than 0.5% of your real users, so you're targeting for the best possible rate of fraud detection without going over 0.5% FRR (all else being equal).

Checks that are graded on a gradient each have their own [sensitivity](./configuration.md#sensitivities), from `Level1` (least strict) to `Level10` (most strict), or `Disabled`.
Rather than setting all of them yourself, you can pick a [verification policy](./configuration.md#verification-policy), which moves several at once.
Any sensitivity you set explicitly wins over the policy, so a policy is a starting point for an evaluation, not a constraint on it.
See [Configure Verify](./configuration.md) for the full set of settings, and [Example configurations](./example-configurations.md) for scenarios to start from.

## Evaluating the verdict

[`verification.verdict`](./verdict.md) returns one of five values for each document: `Accept`, `Reject`, `Review`, `Unverifiable`, and `Retry`.
You don't need to filter your dataset before evaluating—every document falls into one of these buckets.

FRR is the share of real documents that received a `Reject` verdict.
FAR is the share of fake documents that received an `Accept` verdict.

Documents that came back `Unverifiable` or `Retry` were not conclusively classified, so track these separately as your unprocessed rate.

If your [manual review strategy](./configuration.md#manual-review-strategy) is anything other than `Never`, also measure what share of real and fake documents came back as `Review`.
Keep the strategy fixed across runs you intend to compare.
It decides which side of the decision can be diverted into review, so switching it changes your measured rates on its own: `RejectedOnly`, for example, moves borderline documents out of `Reject` and into `Review`, which lowers FRR without any check behaving differently.

## Recording what produced each result

An evaluation is only reproducible if you know exactly what ran.
Store these alongside every result:

- `configurationUsed`, which echoes the effective configuration and mirrors the structure of the request.
  This is also how you find out what a default or a policy resolved to, instead of assuming.
- `verification.failedChecks`, a flat array of dotted check paths such as `checks.extractedDataCheck.matchCheck.dateOfBirthCheck`.
  Aggregating it across the run shows which checks account for most of your false rejections, which is where tuning pays off most.
- `runtime`, which carries `blinkIdVerifyVersion` and `blinkIdVersion` so you can tell results from different releases apart, plus `traceId` and `executionId` for correlating a record with server-side logs when you need to ask about one specific document.

---

[^1]: See how Verify performs on the [DHS IDNet dataset](https://microblink.com/about/newsroom/microblink-identified-as-only-vendor-to-meet-all-performance-thresholds-in-u-s-department-of-homeland-security-identity-verification-evaluation/).


Last updated on Aug 3, 2026
