Evaluate Verify
Before going to production, you may want to evaluate Verify on your own data. The difficulty of this evaluation depends on the type of data that you have.
What to expect
On a global evaluation dataset of hundreds of thousands of real-world images of real and fake documents, Verify achieves highly reliable results with the default configuration.
When controlling for image quality, the false rejection rate (FRR) is low both globally and in the USA.
The evaluation also includes liveness false rejections, but those datasets don't have screen/photocopy presentation attacks. On dedicated liveness datasets, screen/photocopy (and aggregate liveness) false acceptance rate and false rejection rate remain low.1
False acceptance and false rejection
If you have a sample set of real document images, you can measure Verify's false rejection rate.
But if you have only images of real documents, you can't measure the false rejection rate. Inversely, if you only have images of fake documents, you can't measure the false acceptance rate.
The balance between FRR and FAR is important because any system can achieve a perfectly low false rejection score by accepting everything (including fraudulent documents), or a perfectly low false acceptance score by rejecting everything (including real documents).
Be careful when testing with expired documents!
By default, Verify rejects expired documents.
If your dataset contains expired documents you want to treat as genuine, set rejectExpiredDocuments to false.
In theory, you can simply disable all checks and get 0% FRR. Without fake documents, you won't know you're not catching any fraud. Similarly, you might be comparing two products, and if one has 1.5% FRR, while the other one has 1.75% FRR, you might conclude the first one is better. But what if it catches 10 times less fraud then the second one?
We recommend acquiring a quality dataset of synthetic fake documents, like the Department of Homeland Security's IDNet dataset.
Recommended sample size for evaluation
To ensure that the evaluation results accurately reflect BlinkID Verify's performance, it is important to use a large enough and diverse sample set. Evaluations based on a very small number of images (for example, 5 examples per document type) can lead to misleading conclusions due to lack of representative data. We recommend the following sample sizes for reliable testing:
Minimum setup:
- 20 to 30 real document examples per document type
- 20 to 30 fake document examples per document type
Recommended setup:
- around 100 real and 100 fake examples per document type
This way, your evaluation results will more accurately reflect real-world performance.
Trade-offs
FRR and FAR exist in tension.
If you want to increase the number of accepted real documents, the trade-off is that you will also increase the number of accepted fake documents.
If you want to increase the number of rejected fake documents, the trade-off is that you will also increase the number of rejected real documents.
It's important to identify where you want to be on this scale, and tune the solution to get there. This is typically done by fixing one of the two metrics.
For example, you might know you don't want to reject more than 0.5% of your real users, so you're targeting for the best possible rate of fraud detection without going over 0.5% FRR (all else being equal).
Checks that are graded on a gradient each have their own sensitivity, from Level1 (least strict) to Level10 (most strict), or Disabled.
Rather than setting all of them yourself, you can pick a verification policy, which moves several at once.
Any sensitivity you set explicitly wins over the policy, so a policy is a starting point for an evaluation, not a constraint on it.
See Configure Verify for the full set of settings, and Example configurations for scenarios to start from.
Evaluating the verdict
verification.verdict returns one of five values for each document: Accept, Reject, Review, Unverifiable, and Retry.
You don't need to filter your dataset before evaluating—every document falls into one of these buckets.
FRR is the share of real documents that received a Reject verdict.
FAR is the share of fake documents that received an Accept verdict.
Documents that came back Unverifiable or Retry were not conclusively classified, so track these separately as your unprocessed rate.
If your manual review strategy is anything other than Never, also measure what share of real and fake documents came back as Review.
Keep the strategy fixed across runs you intend to compare.
It decides which side of the decision can be diverted into review, so switching it changes your measured rates on its own: RejectedOnly, for example, moves borderline documents out of Reject and into Review, which lowers FRR without any check behaving differently.
Recording what produced each result
An evaluation is only reproducible if you know exactly what ran. Store these alongside every result:
configurationUsed, which echoes the effective configuration and mirrors the structure of the request. This is also how you find out what a default or a policy resolved to, instead of assuming.verification.failedChecks, a flat array of dotted check paths such aschecks.extractedDataCheck.matchCheck.dateOfBirthCheck. Aggregating it across the run shows which checks account for most of your false rejections, which is where tuning pays off most.runtime, which carriesblinkIdVerifyVersionandblinkIdVersionso you can tell results from different releases apart, plustraceIdandexecutionIdfor correlating a record with server-side logs when you need to ask about one specific document.