Environment variables
The on-prem deployment is configured through container environment variables.
Quick reference
A typical deployment only needs the license, and optionally the workflow:
LICENSE_KEY=<license-key>
LICENSE_APPLICATION_ID=<application-id>
DOCVER_WORKFLOW=ExtractAndVerify
License
Both variables are required. A Worker with missing or invalid license configuration never becomes ready. License capabilities are checked separately from the selected workflow.
LICENSE_KEYLICENSE_APPLICATION_ID
Workflow and capacity
DOCVER_WORKFLOW- Selects the runtime mode.
Extractstarts the API and the Workers without the model server, and disables/api/v3/verify.ExtractAndVerifyalso starts the bundled verification models. Any other value stops startup. - default:
ExtractAndVerify
- Selects the runtime mode.
WORKER_COUNT- Number of Worker processes in the container. Must be a positive integer. Also controls the default request concurrency and the number of Workers required for API readiness.
- default:
2
MicroblinkInflightLimit- Maximum number of document requests admitted concurrently by the shared limiter. Non-positive or invalid values fall back to the default.
- default: the value of
WORKER_COUNT
MicroblinkQueueLimit- Maximum number of requests waiting for an admission permit.
0rejects excess requests immediately with a429. A small positive value absorbs short bursts, at the cost of additional memory and latency. - default:
0
- Maximum number of requests waiting for an admission permit.
WORKER_COUNT is capacity, not a throughput guarantee.
Raising it gives the container more simultaneous processing slots, but each Worker also consumes CPU and memory.
Size the container and tune the worker count against a representative workload.
Keep MicroblinkQueueLimit small, or 0, on Kubernetes, and add replicas for sustained load.
For durable buffering, place a shared queue or gateway in front of multiple instances.
A larger bounded internal queue can help a single Docker Compose instance when callers have long enough timeouts, but it should stay small enough to keep overload visible.
Memory and request limits
MicroblinkMemoryGuardPercent- Starts rejecting document requests with a
429when container memory use reaches this percentage of the detected cgroup limit.0disables the guard. - default:
95
- Starts rejecting document requests with a
MicroblinkMemoryGuardReservationFactor- Multiplies the request
Content-Lengthto estimate the memory an admitted request reserves.0disables reservation-based admission, leaving the usage threshold active. - default:
10
- Multiplies the request
MicroblinkMemoryGuardUnknownContentLengthBytes- Base reservation, in bytes, for a request without a
Content-Length. The reservation factor is applied to this value. - default:
20971520
- Base reservation, in bytes, for a request without a
The memory guard depends on a container memory limit. Without one, the API can't calculate the percentage threshold, and the guard is inactive. Always set a Docker or Kubernetes memory limit in production.
Caller, ingress, load balancer, and API timeouts should agree. Raising only one of them doesn't make an end-to-end request live longer if another layer closes the connection first.
Health and shutdown
HEALTH_PORT_BASE- First Worker health-server port.
Worker index
Nlistens onHEALTH_PORT_BASE + N. These ports are used inside the container and don't normally need to be published. - default:
18080
- First Worker health-server port.
Worker index
WorkerHealth__Path- Worker endpoint polled by the API readiness check. Keep the default for the bundled Worker.
- default:
/health/ready
MICROBLINK_SHUTDOWN_GRACE_SECONDS- Seconds the supervisor waits after forwarding a termination signal before killing the remaining service process groups. Must be a non-negative integer. Set the orchestrator's termination grace period above it.
- default:
10
The supervisor treats the API, every Worker, and the model server as one unit. If any managed component exits unexpectedly, the supervisor stops the others and exits, so Docker or Kubernetes can replace the whole instance.
Logging
LOG_FILES_ENABLED- Standard output and error are always used.
Set to
1,TRUE,YES, orONto also mirror component logs under/var/log. The directory must be writable, and should be backed by a volume if you want to keep the logs. - default:
0
- Standard output and error are always used.
Set to
Logging__LogLevel__Default- Default .NET API log level.
- default:
Information
Logging__LogLevel__Microsoft.AspNetCore- ASP.NET Core framework log level.
- default:
Warning
When file logging is enabled, the files are /var/log/api.log, /var/log/worker-N.out.log, /var/log/worker-N.err.log, and, in ExtractAndVerify, /var/log/tf-serving.log.
Container log collection remains the recommended production setup.
API access and CORS
Api__CorsOrigins- Comma-separated allowed browser origins, or
*. Set explicit origins when browser clients call the API directly. - default:
*
- Comma-separated allowed browser origins, or
Api__CorsMethods- Comma-separated allowed CORS methods, or
*. - default:
*
- Comma-separated allowed CORS methods, or
Api__CorsHeaders- Comma-separated allowed CORS request headers, or
*. - default:
*
- Comma-separated allowed CORS request headers, or
CORS is a browser policy. It is not authentication, and it is not network access control. Restrict exposure with the deployment's Service, Ingress, firewall, or gateway.
The API listens on container port 8080.
Change the host, Service, or Ingress port mapping rather than the internal port:
docker run -p 9090:8080 ...
Several image-owned variables assume port 8080, so changing only one internal port variable can disconnect the Workers from the API.
Bundled model-serving controls
These configure the model server and Worker readiness in ExtractAndVerify.
The published image supplies working defaults.
Change them only as part of a tested, deployment-specific tuning exercise.
They don't apply to extraction-only deployments.
TF_SERVING_ENABLE_BATCHING- Enables batching only when the value is exactly
1. - default:
1
- Enables batching only when the value is exactly
TF_SERVING_BATCHING_CONFIG- Batching parameters passed to the model server when batching is enabled. The path must exist inside the container.
- default:
/models/batching.config
TF_NUM_INTRAOP_THREADS- Intra-operation thread count.
- default:
4
TF_NUM_INTEROP_THREADS- Inter-operation thread count.
- default:
2
OMP_NUM_THREADS- OpenMP thread count available to model-serving runtime libraries.
- default:
4
MODEL_SERVING_READINESS_MODE- Worker readiness method. The default verifies every configured model over the same gRPC serving path used for inference.
- default:
tensorflow-signature
MODEL_SERVING_READINESS_TIMEOUT_MS- Per-model readiness request timeout, in milliseconds.
- default:
1000
TF_SERVING_READY_FILE- Optional path used with
tensorflow-ready-filereadiness. The launcher creates the file after the model server reports that model loading has completed. - default: unset
- Optional path used with
The supported readiness modes are:
tensorflow-signature: the image default, and the strongest end-to-end check.tensorflow-rest: checks each model through the REST status API.tensorflow-ready-file: combines a startup marker file with a TCP connection check.- Any other value falls back to a TCP connection check.
Image-owned variables
The image also sets values that define its internal topology and runtime behavior. They're listed here to make the container contract explicit, but they are not tuning controls.
| Variable | Image value | Role |
|---|---|---|
ASPNETCORE_URLS | http://+:8080 | Binds the API inside the container |
API_PORT | 8080 | Internal API port used when the image was built |
API_BASE_URL | http://localhost:8080 | Internal queue endpoint used by Workers |
RESOURCES_PATH | /app/res | Bundled recognition-resource directory |
SCHEMA_PATH | /app/Backend/Schemas/document-verification-request.schema.json | Bundled internal request schema |
DOTNET_gcServer | 1 | Enables server GC for the API process |
DOTNET_GCDynamicAdaptationMode | 1 | Enables dynamic GC adaptation |
DOTNET_GCConserveMemory | 1 | Biases the API GC toward conserving memory |
DOTNET_GCRetainVM | 0 | Returns released GC segments to the OS |
SSL_CERT_FILE | /etc/ssl/certs/ca-certificates.crt | CA bundle for native TLS clients |
CURL_CA_BUNDLE | /etc/ssl/certs/ca-certificates.crt | CA bundle for curl-compatible HTTP clients |
The Worker launcher derives HEALTH_PORT from HEALTH_PORT_BASE, and assigns WORKER_INDEX from 0 through WORKER_COUNT - 1 for each process.
Values set on the container for those two are overwritten.
Changing API_BASE_URL, the bundled resource paths, or the internal serving endpoints turns the deployment into a different topology, and is outside the supported single-image configuration.