Skip to content

Performance & memory

Patronus Ark is built to sit in the request path on ordinary hardware — a laptop CPU, no GPU required. This page explains the design levers that make that possible. For step-by-step tuning, see Tune performance & memory; for the full measurement workflow, see Run the local benchmark.

Where the time goes

The layered design means cost is proportional to how far a scan escalates:

Layer Typical latency (batch 1) Runs on
L1 native input-dependent; large PII/DLP inputs can take milliseconds enabled categories with native detectors
L2 NTDB ~1 ms when configured & cached
L3 transformer tens of ms only promoted requests

Because most traffic is resolved at L1/L2, the expensive transformer runs for only the uncertain minority. This is the single most important performance property of the system.

The memory and latency levers

Static-embedding L2

L2 uses the shared mmBERT tokenizer and a static token-embedding encoder across all categories in a process — no per-token attention pass. The official L3 models are aligned to the same tokenizer and embedding space, and compatible promotions reuse L2's token IDs. One L2 encoder serves seven categories, so L2 is cheap to run and cheap to add categories to without paying for a second tokenization step at promotion. See NTDB format.

Quantized ONNX at L3

L3 transformers use the combined INT8-weight / INT4-embedding ONNX variant by default, so a model bundle is a few hundred MB on disk and resident. Injection, Threat, Sensitive Document, Lion Warden, and the separate Dynamic-PII GLiNER model can select their pinned FP16 graph with PATRONUS_L3_PRECISION=fp16; asset warmup downloads only the selected graph. For Linux x86_64 production, FP16 is the validated choice: use it even though its larger graph can add CPU latency.

Compact tokenizers

Compatible mmBERT-style tokenizers are converted once, on first use, into a compact .mmbpe file. It is generated locally during verified asset downloads and cached warmup when the tokenizer JSON has the supported byte-fallback BPE shape. The generated file stores explicit merge-pair identity, is hash- and version-invalidated, and keeps the canonical Hugging Face tokenizer.json as the conversion source. Runtime classification requires .mmbpe; there is no Hugging Face fallback. The former .kit format is unsupported.

L3 sessions: built on first use, then RAM-resident

An L3 ONNX session is built the first time a scan reaches that model, then held resident in RAM and evicted only after an idle TTL (PATRONUS_L3_TTL_SECS, default 300 s; -1 disables eviction). Models are never hot-swapped per request, so you do not pay repeated load/unload costs under steady traffic — but budget memory for the L3 models your configuration enables, since a hot model stays resident. (The one exception is dynamic-pii: its GLiNER session is warmed eagerly during warmup(), before any request, then follows the same idle-TTL eviction.)

Cost-scheduled worker

The L3 worker schedules promoted jobs by estimated and observed compute cost (an EWMA of real execution time), with a max-wait guard so no request starves, and processes work off the request path so a slow transformer never blocks a fast L1/L2 answer.

Long-text windowing

Long inputs are split into disjoint UTF-8 windows of at most 128 KiB before tokenization. Each window is tokenized once into chunks of at most 254 content tokens. L3 adds BOS/EOS and padding to 256 positions, reusing the exact L2 IDs. The tokenizer input is bounded; retained chunks and results grow with document length. With representative clustering enabled (off by default) near-duplicate windows are grouped by similarity and only cluster representatives run — the rest inherit the representative's verdict — which cuts physical inferences per request on top of the memory bound. See the L3 worker policy.

Unified multi-head L3

Running one coalesced multi-head model (l3_strategy="multi") instead of one model per category lets several promoted categories share a single inference — a large throughput win when multiple model-backed categories are active. Compare both strategies with the local benchmark.

Cache-skip inference for repeats and near-duplicates

The optional model-output cache is the largest latency lever for repetitive traffic: an exact-chunk hit returns stored logits from RAM in microseconds instead of running the transformer (tens of milliseconds), and a close L2-embedding match can propagate a non-safe verdict without any L3 inference. The dynamic-pii GLiNER pipeline caches both exact chunks and known entity spans the same way. The cache is memory-only by default; add a persistent redb file to keep hits across restarts.

Execution backend and threading

The ONNX execution provider is configured via set_execution_backend (CPU, CoreML, CUDA, …). Threading and spin-wait behavior are configured with OnnxRuntimeOptions through set_onnx_runtime_options. Defaults target constrained CPU execution. See the configuration reference.

Measuring on your hardware

Do not trust generic numbers — measure. Every gateway can benchmark itself on the shipped validation samples, reporting latency, throughput, false-positive rate, and per-layer timing for both L3 strategies:

scanner.run_local_benchmark()

See Run the local benchmark.