Skip to main content

Migrate from self-hosted extraction API

Introduction​

The self-hosted extraction API is no longer a separate product.

Both extraction and verification are now bundled in the same image.

This guide covers the extraction API migration only. It doesn't cover verification or new capabilities. See the full documentation and the release notes for that.

OpenAPI specification​

The OpenAPI reference is here and the specification can be found here.

The on-prem and cloud deployments have separate schemas. This guide covers the on-prem API; for the cloud one, see the cloud API reference.

Overview of changes​

  • Two BlinkID document endpoints collapse into one: /api/v3/extract.
  • The request is multipart/form-data only, with binary image parts.
  • Flat request parameters move into a nested configuration object and are renamed.
  • The extracted data moves under extraction.result.

The rest of this guide goes into details.

Endpoints​

The single-side and multi-side endpoints are replaced by one:

POST /api/v3/extract

Which sides you scan is now determined by which image parts you send, not by which URL you call.

  • POST /blinkid-single-side is now POST /api/v3/extract with only the imageFirstSide parameter
  • POST /blinkid-multi-side is now POST /api/v3/extract with imageFirstSide and imageSecondSide parameters
  • POST /barcode is now POST /api/v3/extract-barcode, covered in Barcode-only requests
  • POST /blinkcard is no longer a part of the API and will be present in a separate self-hosted BlinkCard instance.
note

Extraction results are also available in the /api/v3/verify endpoint, with additional verification-related results.

Authentication​

The on-prem API has no per-request authentication.

The only authentication is performed at deployment time, by passing your license as an environment variable.

docker run --rm -p 8080:8080 \
-e LICENSE_KEY={your_license_key} \
-e LICENSE_APPLICATION_ID={your_application_id} \
-e DOCVER_WORKFLOW=Extract \
us-docker.pkg.dev/document-verification-public/on-prem/core:4000.0.0
  • mb-api-key header is removed.
  • X-ClientCustomerId has no replacement: use traceId to correlate a request with your own records.

Request​

Only multipart​

The self-hosted BlinkID API took an application/json body and accepted images either as a URL or as a base64-encoded string:

The on-prem API accepts only multipart/form-data.

curl 'https://{your_host}/api/v3/extract' \
--request POST \
--form 'imageFirstSide=@front.jpg' \
--form 'imageSecondSide=@back.jpg'

Image parameter names​

  • for document images:
    • for multi-side scans, pass both imageFirstSide and imageSecondSide
    • for single-side scans, pass only imageFirstSide
  • for barcode images:
    • use imageBarcode on the /extract endpoint
    • use image on the /extract-barcode

Each image must be JPEG or PNG and at most 25 MiB. Decoded dimensions and aspect ratio are also validated at runtime. See Image requirements.

The old imageFormat request parameter, which let you pick the file type of the returned images is gone. Returned images come back as base64-encoded JPG files.

Settings move into a configuration object​

In the self-hosted extraction API, every setting was a top-level sibling of the images.

In v3, settings go into a single JSON configuration part, structured by module similarly to v8000 scanning settings:

configuration.json
{
"extraction": {
"redactionSettings": {
"globalMode": "ImageOnly"
},
"documentCaptureModuleSettings": {
"faceImageExtractionEnabled": true,
"documentImageReturnEnabled": true,
"secondSideWithNoExtractableDataSkipped": true,
"passportDataPageScanOnly": true
},
"vizModuleSettings": {
"signatureImageExtractionEnabled": false,
"characterValidationEnabled": true
},
"barcodeModuleSettings": {
"barcodeImageReturnEnabled": false
}
},
"imageAssessment": {
"imageQualitySensitivity": "Level5"
}
}

Send it as a form field with a JSON content type:

curl 'https://{your_host}/api/v3/extract' \
--request POST \
--form 'imageFirstSide=@front.jpg' \
--form 'configuration=@configuration.json;type=application/json'

Renamed and removed settings​

Image return​

  • returnFullDocumentImage is now extraction.documentCaptureModuleSettings.documentImageReturnEnabled
  • returnFaceImage is now extraction.documentCaptureModuleSettings.faceImageExtractionEnabled
  • returnSignatureImage is now extraction.vizModuleSettings.signatureImageExtractionEnabled

Anonymization is now redaction​

anonymizationMode is now extraction.redactionSettings.globalMode, and its values change case:

  • NONE is now None
  • IMAGE_ONLY is now ImageOnly
  • RESULT_FIELDS_ONLY is now ResultFieldsOnly
  • FULL_RESULT is now FullResult

The default is still FullResult, so documents that legally require redaction are still redacted if you send no configuration at all.

customDocumentAnonymizationSettings is removed. Custom redaction (anonymizing arbitrary fields from arbitrary documents) will be added in a future release.

Image quality booleans​

The following settings are removed:

  • skipImagesWithBlur
  • skipImagesWithInadequateLightingConditions
  • skipImagesOccludedByHand

They are replaced by imageAssessment.imageQualitySensitivity, which is a general (blur, glare, lighting...) image quality scale. Low sensitivity (for example "Level1") allows blurred and warped images. Higher sensitivity (for example "Level10") rejects blurred and warped images.

Scanning behavior​

  • scanPassportDataPageOnly is now extraction.documentCaptureModuleSettings.passportDataPageScanOnly
  • enableCharacterValidation is now extraction.vizModuleSettings.characterValidationEnabled
  • scanUnsupportedSecondSide is now extraction.documentCaptureModuleSettings.secondSideWithNoExtractableDataSkipped

The last one inverts its meaning along with its name. scanUnsupportedSecondSide: true asked the API to scan a back side it couldn't extract data from; the equivalent is secondSideWithNoExtractableDataSkipped: false.

These three settings exist only on /api/v3/extract. The extraction configuration accepted by /api/v3/verify is a smaller object that doesn't include them.

Other removals​

  • customDocumentRules
  • imageFormat (the API always returns base64-encoded JPG)
  • enableBarcodeScanOnly: use the barcode endpoint for barcode-only requests
  • scanCroppedDocumentImage (both cropped an uncropped images are processed)
  • allowUncertainFirstSideScan
  • maxAllowedMismatchesPerField: cross-field data matching is now a verification concern, tuned with verification.settings.dataMatchSensitivity, and isn't available from the extraction endpoint
  • ageLimit

Response​

Changed response structure​

The self-hosted extraction API returned a JSON object, with data holding the extracted data.

In the on-prem API:

  • extracted data is in extraction.result
  • images are in images
  • timing and diagnostic fields are in runtime

For example:

On-prem API
{
"imageAssessment": {
//...
},
"extraction": {
"result": {
"firstName": { "latin": "JOHN" },
"documentClassInfo": { "country": "croatia" },
// ...
},
},
"images": {
// ...
},
"runtime": {
// ...
}
}

For most fields, data.{field} is now extraction.result.{field}.

However, data.processingStatus is now extraction.processingStatus (it's not part of the result).

  • executionId is now runtime.executionId
  • startTime is now runtime.startedOn
  • finishTime is now runtime.finishedOn
  • traceId is now runtime.traceId

Extracted field names are unchanged​

Fields in the old data object exist in extraction.result under the same name.

One field is added: ethnicity.

Two nested fields change:

  • parentsInfo[].fullName is new
  • driverLicenseDetailedInfo.vehicleClassesInfo[].licenceType has been removed

No more location data for strings​

Many fields had values for possible scripts (latin, cyrillic...) plus the place where that value was read:

Self-hosted extraction API
{
"firstName": {
"latin": "JOHN",
"latinLocation": { "x": 12, "y": 40, "width": 88, "height": 20 },
"cyrillic": null,
"arabic": null,
"greek": null
}
}

The v3 AlphabetString type keeps the four script keys and drops the four *Location keys. You will also never see a null value; if the script wasn't detected on the document, it won't be present in the response.

On-prem API
{
"firstName": {
"latin": "JOHN"
}
}

Per-side results have moved​

  • data.firstSideViz is now extraction.result.subResults.firstSide.viz
  • data.secondSideViz is now extraction.result.subResults.secondSide.viz
  • data.mrzData is now extraction.result.subResults.firstSide.mrz and extraction.result.subResults.secondSide.mrz
  • data.barcode is now extraction.result.subResults.firstSide.barcode and extraction.result.subResults.secondSide.barcode
  • data.firstSideProcessingStatus is now extraction.additionalProcessingInfo.firstSide.processingStatus
  • data.secondSideProcessingStatus is now extraction.additionalProcessingInfo.secondSide.processingStatus
  • data.firstSideAdditionalProcessingInfo is now extraction.additionalProcessingInfo.firstSide
  • data.secondSideAdditionalProcessingInfo is now extraction.additionalProcessingInfo.secondSide

VizResult has the same fields as the old VIZResult, and MrzResult gains six fields: sanitizedDocumentCode, sanitizedDocumentNumber, sanitizedIssuer, sanitizedNationality, sanitizedOpt1, and sanitizedOpt2.

The barcode result nests its raw payload: what was contained in data.barcode is now in extraction.result.subResults.firstSide.barcode.barcodeData and extraction.result.subResults.secondSide.barcode.barcodeData.

Inside barcode data:

  • rawDataBase64 is now rawData
  • barcodeType and uncertain are new

Images​

Images move out of the extracted data and into a top-level images object:

  • data.fullDocumentFirstSideImage.image is now images.firstSideCropped
  • data.fullDocumentSecondSideImage.image is now images.secondSideCropped
  • data.fullDocumentImage.image (single-side) is now images.firstSideCropped
  • data.faceImage.image is now images.face
  • data.signatureImage.image is now images.signature
  • images.barcode is new

Every value in images is a base64-encoded JPEG string. The old faceImage and signatureImage were objects wrapping the image with its location and side; both of those are gone, as described under location data.

Processing status​

processingStatus moves to extraction.processingStatus, and its values change from SCREAMING_SNAKE_CASE to PascalCase.

The conversion is mechanical, so MANDATORY_FIELD_MISSING becomes MandatoryFieldMissing, and every old value carries over under its new spelling except one:

  • DOCUMENT_FILTERED is removed, along with the document filtering that produced it

Three values are new, inherited from BlinkID v8000:

  • MrzDetectionFailed: the MRZ was not detected on the image
  • InputImageNotFocused: the input image was not focused
  • Canceled: processing was canceled

Image analysis is now image assessment​

The old firstSideImageAnalysisResult and secondSideImageAnalysisResult objects reported twelve separate statuses. v3 replaces them with a top-level imageAssessment object holding three checks, each resolving to Pass, Fail, or NotPerformed:

{
"imageAssessment": {
"imageQualityCheck": { "result": "Pass" },
"croppedDocumentCheck": {
"result": "Pass",
"firstSide": { "result": "Pass" },
"secondSide": { "result": "Pass" }
},
"handPresenceCheck": {
"result": "Pass",
"firstSide": { "result": "Pass" },
"secondSide": { "result": "Pass" }
}
}
}

croppedDocumentCheck and handPresenceCheck carry a rolled-up result plus a per-side breakdown. imageQualityCheck has no per-side breakdown: it's a single result for the request. It replaces the old blur/glare/lighting/color statuses.

documentHandOcclusionStatus is now imageAssessment.handPresenceCheck, per side. realIDDetectionStatus is now extraction.additionalProcessingInfo.{side}.realIdDetectionStatus.

New in the response​

configurationUsed echoes back the configuration the request actually ran with, resolved from your configuration part and the defaults for everything you left out.

runtime also reports the engine versions behind the response: blinkIdVersion and blinkIdVerifyVersion, plus blinkIdRecognizerVersion and blinkIdVerifyRecognizerVersion for the underlying model bundles.

Removed from the response​

  • recognitionStatus: the old ResultState (EMPTY, UNCERTAIN, VALID, STAGE_VALID) is gone. Use extraction.processingStatus instead.
  • recognitionMode
  • scanningFirstSideDone: removed. A v3 request is processed as a whole.
  • dataMatchResult: moved to the verification side, as verification.checks.extractedDataCheck.matchCheck. It isn't available from the extraction endpoint.
  • isBelowAgeLimit and age: removed along with the ageLimit request parameter.

Errors​

The old API returned a DefaultResponse with message, traceId, and executionId for every failure.

v3 uses three shapes:

  • Validation failures (400, 415) return a message plus an errors array, where each entry has a code, the path of the offending parameter, and a human-readable reason.
  • Licensing and worker failures (401, 503) return message and reason.
  • Capacity and timeout failures (413, 429, 500, 504) return just message. A 429 also carries a Retry-After header with a suggested delay in seconds, and can return an empty body when the concurrency limiter rather than the memory-pressure limiter rejects the request.

The errors[].path field names the exact configuration path that was rejected, which makes a mistranslated setting easy to spot.

executionId and traceId are returned only for 200 responses.

Barcode-only requests​

The standalone /barcode endpoint becomes POST /api/v3/extract-barcode.

The image part keeps its name, image, and is the only image part the endpoint accepts. Settings go into configuration.extraction.barcodeSettings:

configuration.json
{
"extraction": {
"barcodeSettings": {
"pdf417ScanningEnabled": true,
"qrScanningEnabled": true
}
}
}
curl 'https://{your_host}/api/v3/extract-barcode' \
--request POST \
--form 'image=@barcode.jpg' \
--form 'configuration=@configuration.json;type=application/json'

Renamed symbology toggles​

Every scan{Symbology} boolean becomes {symbology}ScanningEnabled:

  • scanPdf417 is now pdf417ScanningEnabled
  • scanQrCode is now qrScanningEnabled
  • scanUpce is now upceScanningEnabled
  • scanUpca is now upcaScanningEnabled
  • scanCode128 is now code128ScanningEnabled
  • scanCode39 is now code39ScanningEnabled
  • scanEan8 is now ean8ScanningEnabled
  • scanEan13 is now ean13ScanningEnabled
  • scanItf is now itfScanningEnabled

Two symbologies are new: dataMatrixScanningEnabled and aztecScanningEnabled.

barcodeImageReturnEnabled is also new, and returns the barcode image in images.barcode.

Removed barcode settings​

The low-level scanning tuning parameters are gone, with no replacement:

  • autoScaleDetection
  • nullQuietZoneAllowed
  • readCode39AsExtendedData
  • scanInverse
  • scanUncertain
  • slowerThoroughScan

The barcode response is now parsed​

The old endpoint returned only the raw payload, and left parsing to you:

Self-hosted BlinkID
{
"executionId": "01H...",
"data": {
"barcodeType": "PDF417_BARCODE",
"stringData": "@\n\u001e\rANSI 636000...",
"uncertain": false,
"recognitionStatus": "VALID"
}
}

v3 returns the same raw payload plus the parsed fields, using the same BarcodeResult shape as the document endpoint's subResults.{side}.barcode:

On-prem API
{
"extraction": {
"processingStatus": "Success",
"result": {
"barcode": {
"parsed": true,
"firstName": "JOHN",
"lastName": "DOE",
"dateOfBirth": { "day": 1, "month": 2, "year": 1990 },
"addressDetailedInfo": {},
"driverLicenseDetailedInfo": {},
"extendedElements": [],
"barcodeData": {
"barcodeType": "PDF417",
"stringData": "@\n\u001e\rANSI 636000...",
"rawData": "QAoeDUFOU0kgNjM2MDAw...",
"uncertain": false
}
}
}
},
"configurationUsed": {
"extraction": {
"barcodeSettings": {}
}
},
"runtime": {}
}

The barcode response also carries images, which holds images.barcode when you enable barcodeImageReturnEnabled.

The mapping is:

  • data.stringData is now extraction.result.barcode.barcodeData.stringData
  • data.rawData is now extraction.result.barcode.barcodeData.rawData
  • data.barcodeType is now extraction.result.barcode.barcodeData.barcodeType
  • data.uncertain is now extraction.result.barcode.barcodeData.uncertain
  • data.recognitionStatus is removed; use extraction.processingStatus
  • data.detectionPoints is removed, with no replacement

If you were parsing AAMVA payloads yourself, you can now read extraction.result.barcode directly and drop that code. Check parsed first: it tells you whether the raw payload was successfully interpreted.

Deployment​

Pull the new image​

The image is no longer on Dockerhub, but on our own container registry at us-docker.pkg.dev/document-verification-public/on-prem/core.

Before:

docker pull microblink/api:latest

Now:

docker pull us-docker.pkg.dev/document-verification-public/on-prem/core:4000.0.0

The on-prem image is available starting with the 4000.0.0 tag (see note on epoch versioning below).

Licensing​

Your existing licenses continue working with the new image.

Depending on your deployment, set your environment variables in the same way.

LICENSE_KEY=$YOUR_LICENSE_KEY
APPLICATION_ID=$YOUR_APPLICATION_ID

Epoch versioning​

Starting with v3, the on-prem API adopts the epoch versioning scheme for published images.

{EPOCH * 1000 + MAJOR}.MINOR.PATCH

EPOCH: Increment when you make significant or groundbreaking changes.

MAJOR: Increment when you make minor incompatible API changes.

MINOR: Increment when you add functionality in a backwards-compatible manner.

PATCH: Increment when you make backwards-compatible bug fixes.

So, instead of going from v3.22.3 to v4.0.0, the on-prem API goes from v3.22.3 to v4000.0.0.

Subsequent versions will follow the same scheme.

Orchestrate containers​

The deployment model differs from the old extraction API in ways that affect capacity planning.

The on-prem image bundles an API process and a configurable number of workers, and reports its own sizing back in the runtime object of every response: workerCount, workerIndex, inflightLimit, internalQueueSize, cpus, cpuType, and ram.

Two behaviors are worth planning for:

  • The API returns 429 with Retry-After when it's saturated, so a client that previously assumed unlimited concurrency needs retry handling.
  • A request that arrives before the workers are ready returns 503 with a reason, rather than queueing.

runtime.licenseId and runtime.licenseExpiry are also returned on every response, which makes license expiry easy to monitor without a separate endpoint.

There is no single throughput expectation for extraction-only deployments. Capacity depends on the document types, image sizes, request configuration, worker count, and CPU and memory assigned to the container. Benchmark with a workload representative of your production traffic, then use the results to choose WORKER_COUNT and the number of container instances.

Set DOCVER_WORKFLOW=Extract to run only the extraction and barcode endpoints. This mode doesn't start the server-side verification model service, so it generally uses fewer resources than the default ExtractAndVerify mode. Requests to /api/v3/verify are disabled in extraction-only mode.

See also​