Skip to main content

Environment variables

The on-prem deployment is configured through container environment variables.

Quick reference​

A typical deployment only needs the license, and optionally the workflow:

LICENSE_KEY=<license-key>
LICENSE_APPLICATION_ID=<application-id>
DOCVER_WORKFLOW=ExtractAndVerify

License​

Both variables are required. A Worker with missing or invalid license configuration never becomes ready. License capabilities are checked separately from the selected workflow.

  • LICENSE_KEY
  • LICENSE_APPLICATION_ID

Workflow and capacity​

  • DOCVER_WORKFLOW
    • Selects the runtime mode. Extract starts the API and the Workers without the model server, and disables /api/v3/verify. ExtractAndVerify also starts the bundled verification models. Any other value stops startup.
    • default: ExtractAndVerify
  • WORKER_COUNT
    • Number of Worker processes in the container. Must be a positive integer. Also controls the default request concurrency and the number of Workers required for API readiness.
    • default: 2
  • MicroblinkInflightLimit
    • Maximum number of document requests admitted concurrently by the shared limiter. Non-positive or invalid values fall back to the default.
    • default: the value of WORKER_COUNT
  • MicroblinkQueueLimit
    • Maximum number of requests waiting for an admission permit. 0 rejects excess requests immediately with a 429. A small positive value absorbs short bursts, at the cost of additional memory and latency.
    • default: 0

WORKER_COUNT is capacity, not a throughput guarantee. Raising it gives the container more simultaneous processing slots, but each Worker also consumes CPU and memory. Size the container and tune the worker count against a representative workload.

Keep MicroblinkQueueLimit small, or 0, on Kubernetes, and add replicas for sustained load. For durable buffering, place a shared queue or gateway in front of multiple instances. A larger bounded internal queue can help a single Docker Compose instance when callers have long enough timeouts, but it should stay small enough to keep overload visible.

Memory and request limits​

  • MicroblinkMemoryGuardPercent
    • Starts rejecting document requests with a 429 when container memory use reaches this percentage of the detected cgroup limit. 0 disables the guard.
    • default: 95
  • MicroblinkMemoryGuardReservationFactor
    • Multiplies the request Content-Length to estimate the memory an admitted request reserves. 0 disables reservation-based admission, leaving the usage threshold active.
    • default: 10
  • MicroblinkMemoryGuardUnknownContentLengthBytes
    • Base reservation, in bytes, for a request without a Content-Length. The reservation factor is applied to this value.
    • default: 20971520

The memory guard depends on a container memory limit. Without one, the API can't calculate the percentage threshold, and the guard is inactive. Always set a Docker or Kubernetes memory limit in production.

Caller, ingress, load balancer, and API timeouts should agree. Raising only one of them doesn't make an end-to-end request live longer if another layer closes the connection first.

Health and shutdown​

  • HEALTH_PORT_BASE
    • First Worker health-server port. Worker index N listens on HEALTH_PORT_BASE + N. These ports are used inside the container and don't normally need to be published.
    • default: 18080
  • WorkerHealth__Path
    • Worker endpoint polled by the API readiness check. Keep the default for the bundled Worker.
    • default: /health/ready
  • MICROBLINK_SHUTDOWN_GRACE_SECONDS
    • Seconds the supervisor waits after forwarding a termination signal before killing the remaining service process groups. Must be a non-negative integer. Set the orchestrator's termination grace period above it.
    • default: 10

The supervisor treats the API, every Worker, and the model server as one unit. If any managed component exits unexpectedly, the supervisor stops the others and exits, so Docker or Kubernetes can replace the whole instance.

Logging​

  • LOG_FILES_ENABLED
    • Standard output and error are always used. Set to 1, TRUE, YES, or ON to also mirror component logs under /var/log. The directory must be writable, and should be backed by a volume if you want to keep the logs.
    • default: 0
  • Logging__LogLevel__Default
    • Default .NET API log level.
    • default: Information
  • Logging__LogLevel__Microsoft.AspNetCore
    • ASP.NET Core framework log level.
    • default: Warning

When file logging is enabled, the files are /var/log/api.log, /var/log/worker-N.out.log, /var/log/worker-N.err.log, and, in ExtractAndVerify, /var/log/tf-serving.log.

Container log collection remains the recommended production setup.

API access and CORS​

  • Api__CorsOrigins
    • Comma-separated allowed browser origins, or *. Set explicit origins when browser clients call the API directly.
    • default: *
  • Api__CorsMethods
    • Comma-separated allowed CORS methods, or *.
    • default: *
  • Api__CorsHeaders
    • Comma-separated allowed CORS request headers, or *.
    • default: *

CORS is a browser policy. It is not authentication, and it is not network access control. Restrict exposure with the deployment's Service, Ingress, firewall, or gateway.

The API listens on container port 8080. Change the host, Service, or Ingress port mapping rather than the internal port:

docker run -p 9090:8080 ...

Several image-owned variables assume port 8080, so changing only one internal port variable can disconnect the Workers from the API.

Bundled model-serving controls​

These configure the model server and Worker readiness in ExtractAndVerify. The published image supplies working defaults. Change them only as part of a tested, deployment-specific tuning exercise.

They don't apply to extraction-only deployments.

  • TF_SERVING_ENABLE_BATCHING
    • Enables batching only when the value is exactly 1.
    • default: 1
  • TF_SERVING_BATCHING_CONFIG
    • Batching parameters passed to the model server when batching is enabled. The path must exist inside the container.
    • default: /models/batching.config
  • TF_NUM_INTRAOP_THREADS
    • Intra-operation thread count.
    • default: 4
  • TF_NUM_INTEROP_THREADS
    • Inter-operation thread count.
    • default: 2
  • OMP_NUM_THREADS
    • OpenMP thread count available to model-serving runtime libraries.
    • default: 4
  • MODEL_SERVING_READINESS_MODE
    • Worker readiness method. The default verifies every configured model over the same gRPC serving path used for inference.
    • default: tensorflow-signature
  • MODEL_SERVING_READINESS_TIMEOUT_MS
    • Per-model readiness request timeout, in milliseconds.
    • default: 1000
  • TF_SERVING_READY_FILE
    • Optional path used with tensorflow-ready-file readiness. The launcher creates the file after the model server reports that model loading has completed.
    • default: unset

The supported readiness modes are:

  • tensorflow-signature: the image default, and the strongest end-to-end check.
  • tensorflow-rest: checks each model through the REST status API.
  • tensorflow-ready-file: combines a startup marker file with a TCP connection check.
  • Any other value falls back to a TCP connection check.

Image-owned variables​

The image also sets values that define its internal topology and runtime behavior. They're listed here to make the container contract explicit, but they are not tuning controls.

VariableImage valueRole
ASPNETCORE_URLShttp://+:8080Binds the API inside the container
API_PORT8080Internal API port used when the image was built
API_BASE_URLhttp://localhost:8080Internal queue endpoint used by Workers
RESOURCES_PATH/app/resBundled recognition-resource directory
SCHEMA_PATH/app/Backend/Schemas/document-verification-request.schema.jsonBundled internal request schema
DOTNET_gcServer1Enables server GC for the API process
DOTNET_GCDynamicAdaptationMode1Enables dynamic GC adaptation
DOTNET_GCConserveMemory1Biases the API GC toward conserving memory
DOTNET_GCRetainVM0Returns released GC segments to the OS
SSL_CERT_FILE/etc/ssl/certs/ca-certificates.crtCA bundle for native TLS clients
CURL_CA_BUNDLE/etc/ssl/certs/ca-certificates.crtCA bundle for curl-compatible HTTP clients

The Worker launcher derives HEALTH_PORT from HEALTH_PORT_BASE, and assigns WORKER_INDEX from 0 through WORKER_COUNT - 1 for each process. Values set on the container for those two are overwritten.

Changing API_BASE_URL, the bundled resource paths, or the internal serving endpoints turns the deployment into a different topology, and is outside the supported single-image configuration.