---
Source: https://docs.microblink.com/on-prem/migrate-from-sh-extraction-api
Title: Migrate from self-hosted extraction API
Description: Migrate from the self-hosted extraction API (BlinkID) to the on-prem API
---

# Migrate from self-hosted extraction API

## Introduction

The [self-hosted extraction API](/extraction/api/ref) is no longer a separate product.

Both extraction and verification are now bundled in the same image.

This guide covers the extraction API migration only.
It doesn't cover verification or new capabilities.
See the full documentation and the release notes for that.

### OpenAPI specification

The OpenAPI reference is [here](/on-prem/api/ref) and the specification can be found [here](pathname:///on-prem/api/latest/openapi.json).

The on-prem and cloud deployments have separate schemas.
This guide covers the on-prem API; for the cloud one, see the [cloud API reference](/verify/api/ref/v3-cloud).

## Overview of changes

- Two BlinkID document endpoints collapse into one: `/api/v3/extract`.
- The request is `multipart/form-data` only, with binary image parts.
- Flat request parameters move into a nested `configuration` object and are renamed.
- The extracted data moves under `extraction.result`.

The rest of this guide goes into details.

## Endpoints

The single-side and multi-side endpoints are replaced by one:

```
POST /api/v3/extract
```

Which sides you scan is now determined by which image parts you send, not by which URL you call.

- `POST /blinkid-single-side` is now `POST /api/v3/extract` with only the `imageFirstSide` parameter
- `POST /blinkid-multi-side` is now `POST /api/v3/extract` with `imageFirstSide` and `imageSecondSide` parameters
- `POST /barcode` is now `POST /api/v3/extract-barcode`, covered in [Barcode-only requests](#barcode-only-requests)
- `POST /blinkcard` is no longer a part of the API and will be present in a separate self-hosted [BlinkCard](/blinkcard) instance.

:::note

Extraction results are also available in the `/api/v3/verify` endpoint, with additional verification-related results.

:::

## Authentication

The on-prem API has no per-request authentication.

The only authentication is performed at deployment time, by passing your license as an environment variable.

```bash
docker run --rm -p 8080:8080 \
  -e LICENSE_KEY={your_license_key} \
  -e LICENSE_APPLICATION_ID={your_application_id} \
  -e DOCVER_WORKFLOW=Extract \
  us-docker.pkg.dev/document-verification-public/on-prem/core:4000.0.0
```

- `mb-api-key` header is removed.
- `X-ClientCustomerId` has no replacement: use `traceId` to correlate a request with your own records.

## Request

### Only multipart

The self-hosted BlinkID API took an `application/json` body and accepted images either as a URL or as a base64-encoded string:

The on-prem API accepts **only** `multipart/form-data`.

<ApiSample
  method="POST"
  url="https://{your_host}/api/v3/extract"
  formData={[
    { name: "imageFirstSide", fileName: "front.jpg" },
    { name: "imageSecondSide", fileName: "back.jpg" },
  ]}
/>

### Image parameter names


- for document images:
  - for multi-side scans, pass both `imageFirstSide` and `imageSecondSide`
  - for single-side scans, pass only `imageFirstSide`
- for barcode images:
  - use `imageBarcode` on the `/extract` endpoint
  - use `image` on the `/extract-barcode`

Each image must be JPEG or PNG and at most 25 MiB.
Decoded dimensions and aspect ratio are also validated at runtime.
See [Image requirements](/verify/image-requirements).

The old `imageFormat` request parameter, which let you pick the file type of the returned images is gone.
Returned images come back as base64-encoded JPG files.

### Settings move into a configuration object

In the self-hosted extraction API, every setting was a top-level sibling of the images.

In v3, settings go into a single JSON `configuration` part, structured by module similarly to [v8000 scanning settings](/blinkid/migration-v8000#changes-in-scanning-settings):

```json title="configuration.json"
{
  "extraction": {
    "redactionSettings": {
      "globalMode": "ImageOnly"
    },
    "documentCaptureModuleSettings": {
      "faceImageExtractionEnabled": true,
      "documentImageReturnEnabled": true,
      "secondSideWithNoExtractableDataSkipped": true,
      "passportDataPageScanOnly": true
    },
    "vizModuleSettings": {
      "signatureImageExtractionEnabled": false,
      "characterValidationEnabled": true
    },
    "barcodeModuleSettings": {
      "barcodeImageReturnEnabled": false
    }
  },
  "imageAssessment": {
    "imageQualitySensitivity": "Level5"
  }
}
```

Send it as a form field with a JSON content type:

<ApiSample
  method="POST"
  url="https://{your_host}/api/v3/extract"
  formData={[
    { name: "imageFirstSide", fileName: "front.jpg" },
    {
      name: "configuration",
      fileName: "configuration.json",
      contentType: "application/json",
    },
  ]}
/>

### Renamed and removed settings

#### Image return

- `returnFullDocumentImage` is now `extraction.documentCaptureModuleSettings.documentImageReturnEnabled`
- `returnFaceImage` is now `extraction.documentCaptureModuleSettings.faceImageExtractionEnabled`
- `returnSignatureImage` is now `extraction.vizModuleSettings.signatureImageExtractionEnabled`

#### Anonymization is now redaction

`anonymizationMode` is now `extraction.redactionSettings.globalMode`, and its values change case:

- `NONE` is now `None`
- `IMAGE_ONLY` is now `ImageOnly`
- `RESULT_FIELDS_ONLY` is now `ResultFieldsOnly`
- `FULL_RESULT` is now `FullResult`

The default is still `FullResult`, so documents that legally require redaction are still redacted if you send no configuration at all.

`customDocumentAnonymizationSettings` is removed. 
Custom redaction (anonymizing arbitrary fields from arbitrary documents) will be added in a future release.

#### Image quality booleans

The following settings are removed:

- `skipImagesWithBlur`
- `skipImagesWithInadequateLightingConditions`
- `skipImagesOccludedByHand`

They are replaced by `imageAssessment.imageQualitySensitivity`, which is a general (blur, glare, lighting...) image quality scale.
Low sensitivity (for example `"Level1"`) allows blurred and warped images.
Higher sensitivity (for example `"Level10"`) rejects blurred and warped images.

#### Scanning behavior

- `scanPassportDataPageOnly` is now `extraction.documentCaptureModuleSettings.passportDataPageScanOnly`
- `enableCharacterValidation` is now `extraction.vizModuleSettings.characterValidationEnabled`
- `scanUnsupportedSecondSide` is now `extraction.documentCaptureModuleSettings.secondSideWithNoExtractableDataSkipped`

The last one inverts its meaning along with its name.
`scanUnsupportedSecondSide: true` asked the API to scan a back side it couldn't extract data from; the equivalent is `secondSideWithNoExtractableDataSkipped: false`.

These three settings exist only on `/api/v3/extract`.
The extraction configuration accepted by `/api/v3/verify` is a smaller object that doesn't include them.

#### Other removals

- `customDocumentRules`
- `imageFormat` (the API always returns base64-encoded JPG)
- `enableBarcodeScanOnly`: use the [barcode endpoint](#barcode-only-requests) for barcode-only requests
- `scanCroppedDocumentImage` (both cropped an uncropped images are processed)
- `allowUncertainFirstSideScan`
- `maxAllowedMismatchesPerField`: cross-field data matching is now a verification concern, tuned with `verification.settings.dataMatchSensitivity`, and isn't available from the extraction endpoint
- `ageLimit`

## Response

### Changed response structure

The self-hosted extraction API returned a JSON object, with `data` holding the extracted data.

In the on-prem API:

- extracted data is in `extraction.result`
- images are in `images`
- timing and diagnostic fields are in `runtime`

For example:

```json title="On-prem API"
{
  "imageAssessment": {
    //...
  },
  "extraction": {
    "result": {
      "firstName": { "latin": "JOHN" },
      "documentClassInfo": { "country": "croatia" },
      // ...
    },
  },
  "images": {
    // ...
  },
  "runtime": {
    // ...
  }
}
```

For most fields, `data.{field}` is now `extraction.result.{field}`.

However, `data.processingStatus` is now `extraction.processingStatus` (it's not part of the result).

- `executionId` is now `runtime.executionId`
- `startTime` is now `runtime.startedOn`
- `finishTime` is now `runtime.finishedOn`
- `traceId` is now `runtime.traceId`

### Extracted field names are unchanged

Fields in the old `data` object exist in `extraction.result` under the same name.

One field is added: `ethnicity`.

Two nested fields change:

- `parentsInfo[].fullName` is new
- `driverLicenseDetailedInfo.vehicleClassesInfo[].licenceType` has been removed

### No more location data for strings

Many fields had values for possible scripts (latin, cyrillic...) plus the place where that value was read:

```json title="Self-hosted extraction API"
{
  "firstName": {
    "latin": "JOHN",
    "latinLocation": { "x": 12, "y": 40, "width": 88, "height": 20 },
    "cyrillic": null,
    "arabic": null,
    "greek": null
  }
}
```

The v3 `AlphabetString` type keeps the four script keys and drops the four `*Location` keys.
You will also never see a `null` value; if the script wasn't detected on the document, it won't be present in the response.

```json title="On-prem API"
{
  "firstName": {
    "latin": "JOHN"
  }
}
```

### Per-side results have moved

- `data.firstSideViz` is now `extraction.result.subResults.firstSide.viz`
- `data.secondSideViz` is now `extraction.result.subResults.secondSide.viz`
- `data.mrzData` is now `extraction.result.subResults.firstSide.mrz` and `extraction.result.subResults.secondSide.mrz`
- `data.barcode` is now `extraction.result.subResults.firstSide.barcode` and `extraction.result.subResults.secondSide.barcode`
- `data.firstSideProcessingStatus` is now `extraction.additionalProcessingInfo.firstSide.processingStatus`
- `data.secondSideProcessingStatus` is now `extraction.additionalProcessingInfo.secondSide.processingStatus`
- `data.firstSideAdditionalProcessingInfo` is now `extraction.additionalProcessingInfo.firstSide`
- `data.secondSideAdditionalProcessingInfo` is now `extraction.additionalProcessingInfo.secondSide`

`VizResult` has the same fields as the old `VIZResult`, and `MrzResult` gains six fields: `sanitizedDocumentCode`, `sanitizedDocumentNumber`, `sanitizedIssuer`, `sanitizedNationality`, `sanitizedOpt1`, and `sanitizedOpt2`.

The barcode result nests its raw payload: what was contained in `data.barcode` is now in `extraction.result.subResults.firstSide.barcode.barcodeData` and `extraction.result.subResults.secondSide.barcode.barcodeData`.

Inside barcode data:

- `rawDataBase64` is now `rawData`
- `barcodeType` and `uncertain` are new

### Images

Images move out of the extracted data and into a top-level `images` object:

- `data.fullDocumentFirstSideImage.image` is now `images.firstSideCropped`
- `data.fullDocumentSecondSideImage.image` is now `images.secondSideCropped`
- `data.fullDocumentImage.image` (single-side) is now `images.firstSideCropped`
- `data.faceImage.image` is now `images.face`
- `data.signatureImage.image` is now `images.signature`
- `images.barcode` is new

Every value in `images` is a base64-encoded JPEG string.
The old `faceImage` and `signatureImage` were objects wrapping the image with its `location` and `side`; both of those are gone, as described under [location data](#no-more-location-data-for-strings).

### Processing status

`processingStatus` moves to `extraction.processingStatus`, and its values change from `SCREAMING_SNAKE_CASE` to `PascalCase`.

The conversion is mechanical, so `MANDATORY_FIELD_MISSING` becomes `MandatoryFieldMissing`, and every old value carries over under its new spelling except one:

- `DOCUMENT_FILTERED` is removed, along with the document filtering that produced it

Three values are new, inherited from BlinkID v8000:

- `MrzDetectionFailed`: the MRZ was not detected on the image
- `InputImageNotFocused`: the input image was not focused
- `Canceled`: processing was canceled

### Image analysis is now image assessment

The old `firstSideImageAnalysisResult` and `secondSideImageAnalysisResult` objects reported twelve separate statuses.
v3 replaces them with a top-level `imageAssessment` object holding three checks, each resolving to `Pass`, `Fail`, or `NotPerformed`:

```json
{
  "imageAssessment": {
    "imageQualityCheck": { "result": "Pass" },
    "croppedDocumentCheck": {
      "result": "Pass",
      "firstSide": { "result": "Pass" },
      "secondSide": { "result": "Pass" }
    },
    "handPresenceCheck": {
      "result": "Pass",
      "firstSide": { "result": "Pass" },
      "secondSide": { "result": "Pass" }
    }
  }
}
```

`croppedDocumentCheck` and `handPresenceCheck` carry a rolled-up `result` plus a per-side breakdown.
`imageQualityCheck` has no per-side breakdown: it's a single result for the request.
It replaces the old blur/glare/lighting/color statuses.

`documentHandOcclusionStatus` is now `imageAssessment.handPresenceCheck`, per side.
`realIDDetectionStatus` is now `extraction.additionalProcessingInfo.{side}.realIdDetectionStatus`.

### New in the response

`configurationUsed` echoes back the configuration the request actually ran with, resolved from your `configuration` part and the defaults for everything you left out.

`runtime` also reports the engine versions behind the response: `blinkIdVersion` and `blinkIdVerifyVersion`, plus `blinkIdRecognizerVersion` and `blinkIdVerifyRecognizerVersion` for the underlying model bundles.

### Removed from the response

- `recognitionStatus`: the old `ResultState` (`EMPTY`, `UNCERTAIN`, `VALID`, `STAGE_VALID`) is gone. 
  Use `extraction.processingStatus` instead.
- `recognitionMode`
- `scanningFirstSideDone`: removed. A v3 request is processed as a whole.
- `dataMatchResult`: moved to the verification side, as `verification.checks.extractedDataCheck.matchCheck`. 
  It isn't available from the extraction endpoint.
- `isBelowAgeLimit` and `age`: removed along with the `ageLimit` request parameter.

### Errors

The old API returned a `DefaultResponse` with `message`, `traceId`, and `executionId` for every failure.

v3 uses three shapes:

- Validation failures (`400`, `415`) return a `message` plus an `errors` array, where each entry has a `code`, the `path` of the offending parameter, and a human-readable `reason`. 
- Licensing and worker failures (`401`, `503`) return `message` and `reason`.
- Capacity and timeout failures (`413`, `429`, `500`, `504`) return just `message`. 
  A `429` also carries a `Retry-After` header with a suggested delay in seconds, and can return an empty body when the concurrency limiter rather than the memory-pressure limiter rejects the request.

The `errors[].path` field names the exact configuration path that was rejected, which makes a mistranslated setting easy to spot.

`executionId` and `traceId` are returned only for 200 responses.

## Barcode-only requests

The standalone `/barcode` endpoint becomes `POST /api/v3/extract-barcode`.

The image part keeps its name, `image`, and is the only image part the endpoint accepts.
Settings go into `configuration.extraction.barcodeSettings`:

```json title="configuration.json"
{
  "extraction": {
    "barcodeSettings": {
      "pdf417ScanningEnabled": true,
      "qrScanningEnabled": true
    }
  }
}
```

<ApiSample
  method="POST"
  url="https://{your_host}/api/v3/extract-barcode"
  formData={[
    { name: "image", fileName: "barcode.jpg" },
    {
      name: "configuration",
      fileName: "configuration.json",
      contentType: "application/json",
    },
  ]}
/>

### Renamed symbology toggles

Every `scan{Symbology}` boolean becomes `{symbology}ScanningEnabled`:

- `scanPdf417` is now `pdf417ScanningEnabled`
- `scanQrCode` is now `qrScanningEnabled`
- `scanUpce` is now `upceScanningEnabled`
- `scanUpca` is now `upcaScanningEnabled`
- `scanCode128` is now `code128ScanningEnabled`
- `scanCode39` is now `code39ScanningEnabled`
- `scanEan8` is now `ean8ScanningEnabled`
- `scanEan13` is now `ean13ScanningEnabled`
- `scanItf` is now `itfScanningEnabled`

Two symbologies are new: `dataMatrixScanningEnabled` and `aztecScanningEnabled`.

`barcodeImageReturnEnabled` is also new, and returns the barcode image in `images.barcode`.

### Removed barcode settings

The low-level scanning tuning parameters are gone, with no replacement:

- `autoScaleDetection`
- `nullQuietZoneAllowed`
- `readCode39AsExtendedData`
- `scanInverse`
- `scanUncertain`
- `slowerThoroughScan`

### The barcode response is now parsed

The old endpoint returned only the raw payload, and left parsing to you:

```json title="Self-hosted BlinkID"
{
  "executionId": "01H...",
  "data": {
    "barcodeType": "PDF417_BARCODE",
    "stringData": "@\n\u001e\rANSI 636000...",
    "uncertain": false,
    "recognitionStatus": "VALID"
  }
}
```

v3 returns the same raw payload plus the parsed fields, using the same `BarcodeResult` shape as the document endpoint's `subResults.{side}.barcode`:

```json title="On-prem API"
{
  "extraction": {
    "processingStatus": "Success",
    "result": {
      "barcode": {
        "parsed": true,
        "firstName": "JOHN",
        "lastName": "DOE",
        "dateOfBirth": { "day": 1, "month": 2, "year": 1990 },
        "addressDetailedInfo": {},
        "driverLicenseDetailedInfo": {},
        "extendedElements": [],
        "barcodeData": {
          "barcodeType": "PDF417",
          "stringData": "@\n\u001e\rANSI 636000...",
          "rawData": "QAoeDUFOU0kgNjM2MDAw...",
          "uncertain": false
        }
      }
    }
  },
  "configurationUsed": {
    "extraction": {
      "barcodeSettings": {}
    }
  },
  "runtime": {}
}
```

The barcode response also carries `images`, which holds `images.barcode` when you enable `barcodeImageReturnEnabled`.

The mapping is:

- `data.stringData` is now `extraction.result.barcode.barcodeData.stringData`
- `data.rawData` is now `extraction.result.barcode.barcodeData.rawData`
- `data.barcodeType` is now `extraction.result.barcode.barcodeData.barcodeType`
- `data.uncertain` is now `extraction.result.barcode.barcodeData.uncertain`
- `data.recognitionStatus` is removed; use `extraction.processingStatus`
- `data.detectionPoints` is removed, with no replacement

If you were parsing AAMVA payloads yourself, you can now read `extraction.result.barcode` directly and drop that code.
Check `parsed` first: it tells you whether the raw payload was successfully interpreted.

## Deployment

### Pull the new image

The image is no longer on Dockerhub, but on our own container registry at `us-docker.pkg.dev/document-verification-public/on-prem/core`.

Before:

```bash
docker pull microblink/api:latest
```

Now:

```bash
docker pull us-docker.pkg.dev/document-verification-public/on-prem/core:4000.0.0
```

The on-prem image is available starting with the `4000.0.0` tag (see note on [epoch versioning](#epoch-versioning) below).

### Licensing

Your existing licenses continue working with the new image.

Depending on your deployment, set your environment variables in the same way.

```bash
LICENSE_KEY=$YOUR_LICENSE_KEY
APPLICATION_ID=$YOUR_APPLICATION_ID
```

### Epoch versioning

Starting with v3, the on-prem API adopts the [epoch versioning scheme](https://antfu.me/posts/epoch-semver#epoch-semantic-versioning) for published images.

> `{EPOCH * 1000 + MAJOR}.MINOR.PATCH`
>
> EPOCH: Increment when you make significant or groundbreaking changes.
>
> MAJOR: Increment when you make minor incompatible API changes.
>
> MINOR: Increment when you add functionality in a backwards-compatible manner.
>
> PATCH: Increment when you make backwards-compatible bug fixes.

So, instead of going from v3.22.3 to v4.0.0, **the on-prem API goes from v3.22.3 to v4000.0.0.**

Subsequent versions will follow the same scheme.

### Orchestrate containers

The deployment model differs from the old extraction API in ways that affect capacity planning.

The on-prem image bundles an API process and a configurable number of workers, and reports its own sizing back in the `runtime` object of every response: `workerCount`, `workerIndex`, `inflightLimit`, `internalQueueSize`, `cpus`, `cpuType`, and `ram`.

Two behaviors are worth planning for:

- The API returns `429` with `Retry-After` when it's saturated, so a client that previously assumed unlimited concurrency needs retry handling.
- A request that arrives before the workers are ready returns `503` with a `reason`, rather than queueing.

`runtime.licenseId` and `runtime.licenseExpiry` are also returned on every response, which makes license expiry easy to monitor without a separate endpoint.

There is no single throughput expectation for extraction-only deployments.
Capacity depends on the document types, image sizes, request configuration, worker count, and CPU and memory assigned to the container.
Benchmark with a workload representative of your production traffic, then use the results to choose `WORKER_COUNT` and the number of container instances.

Set `DOCVER_WORKFLOW=Extract` to run only the extraction and barcode endpoints.
This mode doesn't start the server-side verification model service, so it generally uses fewer resources than the default `ExtractAndVerify` mode.
Requests to `/api/v3/verify` are disabled in extraction-only mode.

## See also

- [Migrate to Verify v3](/verify/migrate-v3): if you're also moving from the v2 verification API.
- [BlinkID v8000 migration guide](/blinkid/migration-v8000)
- [Configuration](/verify/configuration): the full configuration reference.
- [Quick start](/on-prem/quick-start): running the on-prem container.
- [On-prem API reference](/on-prem/api)
- [Self-hosted extraction API reference](/extraction/api/ref): the API you're migrating away from.


Last updated on Sep 11, 2026
