Migrate from self-hosted extraction API
Introduction
The self-hosted extraction API is no longer a separate product.
Both extraction and verification are now bundled in the same image.
This guide covers the extraction API migration only. It doesn't cover verification or new capabilities. See the full documentation and the release notes for that.
OpenAPI specification
The OpenAPI reference is here and the specification can be found here.
The on-prem and cloud deployments have separate schemas. This guide covers the on-prem API; for the cloud one, see the cloud API reference.
Overview of changes
- Two BlinkID document endpoints collapse into one:
/api/v3/extract. - The request is
multipart/form-dataonly, with binary image parts. - Flat request parameters move into a nested
configurationobject and are renamed. - The extracted data moves under
extraction.result.
The rest of this guide goes into details.
Endpoints
The single-side and multi-side endpoints are replaced by one:
POST /api/v3/extract
Which sides you scan is now determined by which image parts you send, not by which URL you call.
POST /blinkid-single-sideis nowPOST /api/v3/extractwith only theimageFirstSideparameterPOST /blinkid-multi-sideis nowPOST /api/v3/extractwithimageFirstSideandimageSecondSideparametersPOST /barcodeis nowPOST /api/v3/extract-barcode, covered in Barcode-only requestsPOST /blinkcardis no longer a part of the API and will be present in a separate self-hosted BlinkCard instance.
Extraction results are also available in the /api/v3/verify endpoint, with additional verification-related results.
Authentication
The on-prem API has no per-request authentication.
The only authentication is performed at deployment time, by passing your license as an environment variable.
docker run --rm -p 8080:8080 \
-e LICENSE_KEY={your_license_key} \
-e LICENSE_APPLICATION_ID={your_application_id} \
-e DOCVER_WORKFLOW=Extract \
us-docker.pkg.dev/document-verification-public/on-prem/core:4000.0.0
mb-api-keyheader is removed.X-ClientCustomerIdhas no replacement: usetraceIdto correlate a request with your own records.
Request
Only multipart
The self-hosted BlinkID API took an application/json body and accepted images either as a URL or as a base64-encoded string:
The on-prem API accepts only multipart/form-data.
- cURL
- JavaScript
- Python
- Go
curl 'https://{your_host}/api/v3/extract' \
--request POST \
--form 'imageFirstSide=@front.jpg' \
--form 'imageSecondSide=@back.jpg'
const formData = new FormData()
formData.append('imageFirstSide', new Blob([]), 'front.jpg')
formData.append('imageSecondSide', new Blob([]), 'back.jpg')
fetch('https://{your_host}/api/v3/extract', {
method: 'POST',
body: formData
})
requests.post("https://{your_host}/api/v3/extract",
files=[
("imageFirstSide", open("front.jpg", "rb")),
("imageSecondSide", open("back.jpg", "rb"))
]
)
package main
import (
"bytes"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
)
func main() {
requestUrl := "https://{your_host}/api/v3/extract"
payload := &bytes.Buffer{}
writer := multipart.NewWriter(payload)
part, _ := writer.CreateFormFile("imageFirstSide", "front.jpg")
f, _ := os.Open("front.jpg")
defer f.Close()
_, _ = io.Copy(part, f)
part, _ = writer.CreateFormFile("imageSecondSide", "back.jpg")
f, _ = os.Open("back.jpg")
defer f.Close()
_, _ = io.Copy(part, f)
writer.Close()
req, _ := http.NewRequest("POST", requestUrl, payload)
req.Header.Set("Content-Type", writer.FormDataContentType())
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
body, _ := io.ReadAll(res.Body)
fmt.Println(res)
fmt.Println(string(body))
}
Image parameter names
- for document images:
- for multi-side scans, pass both
imageFirstSideandimageSecondSide - for single-side scans, pass only
imageFirstSide
- for multi-side scans, pass both
- for barcode images:
- use
imageBarcodeon the/extractendpoint - use
imageon the/extract-barcode
- use
Each image must be JPEG or PNG and at most 25 MiB. Decoded dimensions and aspect ratio are also validated at runtime. See Image requirements.
The old imageFormat request parameter, which let you pick the file type of the returned images is gone.
Returned images come back as base64-encoded JPG files.
Settings move into a configuration object
In the self-hosted extraction API, every setting was a top-level sibling of the images.
In v3, settings go into a single JSON configuration part, structured by module similarly to v8000 scanning settings:
{
"extraction": {
"redactionSettings": {
"globalMode": "ImageOnly"
},
"documentCaptureModuleSettings": {
"faceImageExtractionEnabled": true,
"documentImageReturnEnabled": true,
"secondSideWithNoExtractableDataSkipped": true,
"passportDataPageScanOnly": true
},
"vizModuleSettings": {
"signatureImageExtractionEnabled": false,
"characterValidationEnabled": true
},
"barcodeModuleSettings": {
"barcodeImageReturnEnabled": false
}
},
"imageAssessment": {
"imageQualitySensitivity": "Level5"
}
}
Send it as a form field with a JSON content type:
- cURL
- JavaScript
- Python
- Go
curl 'https://{your_host}/api/v3/extract' \
--request POST \
--form 'imageFirstSide=@front.jpg' \
--form 'configuration=@configuration.json;type=application/json'
const formData = new FormData()
formData.append('imageFirstSide', new Blob([]), 'front.jpg')
formData.append('configuration', new Blob([]), 'configuration.json')
fetch('https://{your_host}/api/v3/extract', {
method: 'POST',
body: formData
})
requests.post("https://{your_host}/api/v3/extract",
files=[
("imageFirstSide", open("front.jpg", "rb")),
("configuration", ("configuration.json", open("configuration.json", "rb"), "application/json"))
]
)
package main
import (
"bytes"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
)
func main() {
requestUrl := "https://{your_host}/api/v3/extract"
payload := &bytes.Buffer{}
writer := multipart.NewWriter(payload)
part, _ := writer.CreateFormFile("imageFirstSide", "front.jpg")
f, _ := os.Open("front.jpg")
defer f.Close()
_, _ = io.Copy(part, f)
part, _ = writer.CreateFormFile("configuration", "configuration.json")
f, _ = os.Open("configuration.json")
defer f.Close()
_, _ = io.Copy(part, f)
writer.Close()
req, _ := http.NewRequest("POST", requestUrl, payload)
req.Header.Set("Content-Type", writer.FormDataContentType())
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
body, _ := io.ReadAll(res.Body)
fmt.Println(res)
fmt.Println(string(body))
}
Renamed and removed settings
Image return
returnFullDocumentImageis nowextraction.documentCaptureModuleSettings.documentImageReturnEnabledreturnFaceImageis nowextraction.documentCaptureModuleSettings.faceImageExtractionEnabledreturnSignatureImageis nowextraction.vizModuleSettings.signatureImageExtractionEnabled
Anonymization is now redaction
anonymizationMode is now extraction.redactionSettings.globalMode, and its values change case:
NONEis nowNoneIMAGE_ONLYis nowImageOnlyRESULT_FIELDS_ONLYis nowResultFieldsOnlyFULL_RESULTis nowFullResult
The default is still FullResult, so documents that legally require redaction are still redacted if you send no configuration at all.
customDocumentAnonymizationSettings is removed.
Custom redaction (anonymizing arbitrary fields from arbitrary documents) will be added in a future release.
Image quality booleans
The following settings are removed:
skipImagesWithBlurskipImagesWithInadequateLightingConditionsskipImagesOccludedByHand
They are replaced by imageAssessment.imageQualitySensitivity, which is a general (blur, glare, lighting...) image quality scale.
Low sensitivity (for example "Level1") allows blurred and warped images.
Higher sensitivity (for example "Level10") rejects blurred and warped images.
Scanning behavior
scanPassportDataPageOnlyis nowextraction.documentCaptureModuleSettings.passportDataPageScanOnlyenableCharacterValidationis nowextraction.vizModuleSettings.characterValidationEnabledscanUnsupportedSecondSideis nowextraction.documentCaptureModuleSettings.secondSideWithNoExtractableDataSkipped
The last one inverts its meaning along with its name.
scanUnsupportedSecondSide: true asked the API to scan a back side it couldn't extract data from; the equivalent is secondSideWithNoExtractableDataSkipped: false.
These three settings exist only on /api/v3/extract.
The extraction configuration accepted by /api/v3/verify is a smaller object that doesn't include them.
Other removals
customDocumentRulesimageFormat(the API always returns base64-encoded JPG)enableBarcodeScanOnly: use the barcode endpoint for barcode-only requestsscanCroppedDocumentImage(both cropped an uncropped images are processed)allowUncertainFirstSideScanmaxAllowedMismatchesPerField: cross-field data matching is now a verification concern, tuned withverification.settings.dataMatchSensitivity, and isn't available from the extraction endpointageLimit
Response
Changed response structure
The self-hosted extraction API returned a JSON object, with data holding the extracted data.
In the on-prem API:
- extracted data is in
extraction.result - images are in
images - timing and diagnostic fields are in
runtime
For example:
{
"imageAssessment": {
//...
},
"extraction": {
"result": {
"firstName": { "latin": "JOHN" },
"documentClassInfo": { "country": "croatia" },
// ...
},
},
"images": {
// ...
},
"runtime": {
// ...
}
}
For most fields, data.{field} is now extraction.result.{field}.
However, data.processingStatus is now extraction.processingStatus (it's not part of the result).
executionIdis nowruntime.executionIdstartTimeis nowruntime.startedOnfinishTimeis nowruntime.finishedOntraceIdis nowruntime.traceId
Extracted field names are unchanged
Fields in the old data object exist in extraction.result under the same name.
One field is added: ethnicity.
Two nested fields change:
parentsInfo[].fullNameis newdriverLicenseDetailedInfo.vehicleClassesInfo[].licenceTypehas been removed
No more location data for strings
Many fields had values for possible scripts (latin, cyrillic...) plus the place where that value was read:
{
"firstName": {
"latin": "JOHN",
"latinLocation": { "x": 12, "y": 40, "width": 88, "height": 20 },
"cyrillic": null,
"arabic": null,
"greek": null
}
}
The v3 AlphabetString type keeps the four script keys and drops the four *Location keys.
You will also never see a null value; if the script wasn't detected on the document, it won't be present in the response.
{
"firstName": {
"latin": "JOHN"
}
}
Per-side results have moved
data.firstSideVizis nowextraction.result.subResults.firstSide.vizdata.secondSideVizis nowextraction.result.subResults.secondSide.vizdata.mrzDatais nowextraction.result.subResults.firstSide.mrzandextraction.result.subResults.secondSide.mrzdata.barcodeis nowextraction.result.subResults.firstSide.barcodeandextraction.result.subResults.secondSide.barcodedata.firstSideProcessingStatusis nowextraction.additionalProcessingInfo.firstSide.processingStatusdata.secondSideProcessingStatusis nowextraction.additionalProcessingInfo.secondSide.processingStatusdata.firstSideAdditionalProcessingInfois nowextraction.additionalProcessingInfo.firstSidedata.secondSideAdditionalProcessingInfois nowextraction.additionalProcessingInfo.secondSide
VizResult has the same fields as the old VIZResult, and MrzResult gains six fields: sanitizedDocumentCode, sanitizedDocumentNumber, sanitizedIssuer, sanitizedNationality, sanitizedOpt1, and sanitizedOpt2.
The barcode result nests its raw payload: what was contained in data.barcode is now in extraction.result.subResults.firstSide.barcode.barcodeData and extraction.result.subResults.secondSide.barcode.barcodeData.
Inside barcode data:
rawDataBase64is nowrawDatabarcodeTypeanduncertainare new
Images
Images move out of the extracted data and into a top-level images object:
data.fullDocumentFirstSideImage.imageis nowimages.firstSideCroppeddata.fullDocumentSecondSideImage.imageis nowimages.secondSideCroppeddata.fullDocumentImage.image(single-side) is nowimages.firstSideCroppeddata.faceImage.imageis nowimages.facedata.signatureImage.imageis nowimages.signatureimages.barcodeis new
Every value in images is a base64-encoded JPEG string.
The old faceImage and signatureImage were objects wrapping the image with its location and side; both of those are gone, as described under location data.
Processing status
processingStatus moves to extraction.processingStatus, and its values change from SCREAMING_SNAKE_CASE to PascalCase.
The conversion is mechanical, so MANDATORY_FIELD_MISSING becomes MandatoryFieldMissing, and every old value carries over under its new spelling except one:
DOCUMENT_FILTEREDis removed, along with the document filtering that produced it
Three values are new, inherited from BlinkID v8000:
MrzDetectionFailed: the MRZ was not detected on the imageInputImageNotFocused: the input image was not focusedCanceled: processing was canceled
Image analysis is now image assessment
The old firstSideImageAnalysisResult and secondSideImageAnalysisResult objects reported twelve separate statuses.
v3 replaces them with a top-level imageAssessment object holding three checks, each resolving to Pass, Fail, or NotPerformed:
{
"imageAssessment": {
"imageQualityCheck": { "result": "Pass" },
"croppedDocumentCheck": {
"result": "Pass",
"firstSide": { "result": "Pass" },
"secondSide": { "result": "Pass" }
},
"handPresenceCheck": {
"result": "Pass",
"firstSide": { "result": "Pass" },
"secondSide": { "result": "Pass" }
}
}
}
croppedDocumentCheck and handPresenceCheck carry a rolled-up result plus a per-side breakdown.
imageQualityCheck has no per-side breakdown: it's a single result for the request.
It replaces the old blur/glare/lighting/color statuses.
documentHandOcclusionStatus is now imageAssessment.handPresenceCheck, per side.
realIDDetectionStatus is now extraction.additionalProcessingInfo.{side}.realIdDetectionStatus.
New in the response
configurationUsed echoes back the configuration the request actually ran with, resolved from your configuration part and the defaults for everything you left out.
runtime also reports the engine versions behind the response: blinkIdVersion and blinkIdVerifyVersion, plus blinkIdRecognizerVersion and blinkIdVerifyRecognizerVersion for the underlying model bundles.
Removed from the response
recognitionStatus: the oldResultState(EMPTY,UNCERTAIN,VALID,STAGE_VALID) is gone. Useextraction.processingStatusinstead.recognitionModescanningFirstSideDone: removed. A v3 request is processed as a whole.dataMatchResult: moved to the verification side, asverification.checks.extractedDataCheck.matchCheck. It isn't available from the extraction endpoint.isBelowAgeLimitandage: removed along with theageLimitrequest parameter.
Errors
The old API returned a DefaultResponse with message, traceId, and executionId for every failure.
v3 uses three shapes:
- Validation failures (
400,415) return amessageplus anerrorsarray, where each entry has acode, thepathof the offending parameter, and a human-readablereason. - Licensing and worker failures (
401,503) returnmessageandreason. - Capacity and timeout failures (
413,429,500,504) return justmessage. A429also carries aRetry-Afterheader with a suggested delay in seconds, and can return an empty body when the concurrency limiter rather than the memory-pressure limiter rejects the request.
The errors[].path field names the exact configuration path that was rejected, which makes a mistranslated setting easy to spot.
executionId and traceId are returned only for 200 responses.
Barcode-only requests
The standalone /barcode endpoint becomes POST /api/v3/extract-barcode.
The image part keeps its name, image, and is the only image part the endpoint accepts.
Settings go into configuration.extraction.barcodeSettings:
{
"extraction": {
"barcodeSettings": {
"pdf417ScanningEnabled": true,
"qrScanningEnabled": true
}
}
}
- cURL
- JavaScript
- Python
- Go
curl 'https://{your_host}/api/v3/extract-barcode' \
--request POST \
--form 'image=@barcode.jpg' \
--form 'configuration=@configuration.json;type=application/json'
const formData = new FormData()
formData.append('image', new Blob([]), 'barcode.jpg')
formData.append('configuration', new Blob([]), 'configuration.json')
fetch('https://{your_host}/api/v3/extract-barcode', {
method: 'POST',
body: formData
})
requests.post(
"https://{your_host}/api/v3/extract-barcode",
files=[
("image", open("barcode.jpg", "rb")),
("configuration", ("configuration.json", open("configuration.json", "rb"), "application/json"))
]
)
package main
import (
"bytes"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
)
func main() {
requestUrl := "https://{your_host}/api/v3/extract-barcode"
payload := &bytes.Buffer{}
writer := multipart.NewWriter(payload)
part, _ := writer.CreateFormFile("image", "barcode.jpg")
f, _ := os.Open("barcode.jpg")
defer f.Close()
_, _ = io.Copy(part, f)
part, _ = writer.CreateFormFile("configuration", "configuration.json")
f, _ = os.Open("configuration.json")
defer f.Close()
_, _ = io.Copy(part, f)
writer.Close()
req, _ := http.NewRequest("POST", requestUrl, payload)
req.Header.Set("Content-Type", writer.FormDataContentType())
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
body, _ := io.ReadAll(res.Body)
fmt.Println(res)
fmt.Println(string(body))
}
Renamed symbology toggles
Every scan{Symbology} boolean becomes {symbology}ScanningEnabled:
scanPdf417is nowpdf417ScanningEnabledscanQrCodeis nowqrScanningEnabledscanUpceis nowupceScanningEnabledscanUpcais nowupcaScanningEnabledscanCode128is nowcode128ScanningEnabledscanCode39is nowcode39ScanningEnabledscanEan8is nowean8ScanningEnabledscanEan13is nowean13ScanningEnabledscanItfis nowitfScanningEnabled
Two symbologies are new: dataMatrixScanningEnabled and aztecScanningEnabled.
barcodeImageReturnEnabled is also new, and returns the barcode image in images.barcode.
Removed barcode settings
The low-level scanning tuning parameters are gone, with no replacement:
autoScaleDetectionnullQuietZoneAllowedreadCode39AsExtendedDatascanInversescanUncertainslowerThoroughScan
The barcode response is now parsed
The old endpoint returned only the raw payload, and left parsing to you:
{
"executionId": "01H...",
"data": {
"barcodeType": "PDF417_BARCODE",
"stringData": "@\n\u001e\rANSI 636000...",
"uncertain": false,
"recognitionStatus": "VALID"
}
}
v3 returns the same raw payload plus the parsed fields, using the same BarcodeResult shape as the document endpoint's subResults.{side}.barcode:
{
"extraction": {
"processingStatus": "Success",
"result": {
"barcode": {
"parsed": true,
"firstName": "JOHN",
"lastName": "DOE",
"dateOfBirth": { "day": 1, "month": 2, "year": 1990 },
"addressDetailedInfo": {},
"driverLicenseDetailedInfo": {},
"extendedElements": [],
"barcodeData": {
"barcodeType": "PDF417",
"stringData": "@\n\u001e\rANSI 636000...",
"rawData": "QAoeDUFOU0kgNjM2MDAw...",
"uncertain": false
}
}
}
},
"configurationUsed": {
"extraction": {
"barcodeSettings": {}
}
},
"runtime": {}
}
The barcode response also carries images, which holds images.barcode when you enable barcodeImageReturnEnabled.
The mapping is:
data.stringDatais nowextraction.result.barcode.barcodeData.stringDatadata.rawDatais nowextraction.result.barcode.barcodeData.rawDatadata.barcodeTypeis nowextraction.result.barcode.barcodeData.barcodeTypedata.uncertainis nowextraction.result.barcode.barcodeData.uncertaindata.recognitionStatusis removed; useextraction.processingStatusdata.detectionPointsis removed, with no replacement
If you were parsing AAMVA payloads yourself, you can now read extraction.result.barcode directly and drop that code.
Check parsed first: it tells you whether the raw payload was successfully interpreted.
Deployment
Pull the new image
The image is no longer on Dockerhub, but on our own container registry at us-docker.pkg.dev/document-verification-public/on-prem/core.
Before:
docker pull microblink/api:latest
Now:
docker pull us-docker.pkg.dev/document-verification-public/on-prem/core:4000.0.0
The on-prem image is available starting with the 4000.0.0 tag (see note on epoch versioning below).
Licensing
Your existing licenses continue working with the new image.
Depending on your deployment, set your environment variables in the same way.
LICENSE_KEY=$YOUR_LICENSE_KEY
APPLICATION_ID=$YOUR_APPLICATION_ID
Epoch versioning
Starting with v3, the on-prem API adopts the epoch versioning scheme for published images.
{EPOCH * 1000 + MAJOR}.MINOR.PATCHEPOCH: Increment when you make significant or groundbreaking changes.
MAJOR: Increment when you make minor incompatible API changes.
MINOR: Increment when you add functionality in a backwards-compatible manner.
PATCH: Increment when you make backwards-compatible bug fixes.
So, instead of going from v3.22.3 to v4.0.0, the on-prem API goes from v3.22.3 to v4000.0.0.
Subsequent versions will follow the same scheme.
Orchestrate containers
The deployment model differs from the old extraction API in ways that affect capacity planning.
The on-prem image bundles an API process and a configurable number of workers, and reports its own sizing back in the runtime object of every response: workerCount, workerIndex, inflightLimit, internalQueueSize, cpus, cpuType, and ram.
Two behaviors are worth planning for:
- The API returns
429withRetry-Afterwhen it's saturated, so a client that previously assumed unlimited concurrency needs retry handling. - A request that arrives before the workers are ready returns
503with areason, rather than queueing.
runtime.licenseId and runtime.licenseExpiry are also returned on every response, which makes license expiry easy to monitor without a separate endpoint.
There is no single throughput expectation for extraction-only deployments.
Capacity depends on the document types, image sizes, request configuration, worker count, and CPU and memory assigned to the container.
Benchmark with a workload representative of your production traffic, then use the results to choose WORKER_COUNT and the number of container instances.
Set DOCVER_WORKFLOW=Extract to run only the extraction and barcode endpoints.
This mode doesn't start the server-side verification model service, so it generally uses fewer resources than the default ExtractAndVerify mode.
Requests to /api/v3/verify are disabled in extraction-only mode.
See also
- Migrate to Verify v3: if you're also moving from the v2 verification API.
- BlinkID v8000 migration guide
- Configuration: the full configuration reference.
- Quick start: running the on-prem container.
- On-prem API reference
- Self-hosted extraction API reference: the API you're migrating away from.