Layered scanning (L1 / L2 / L3)¶
Patronus Ark scans each category with up to three layers. The layers escalate: cheap, always-available checks run first, and expensive models run only for the traffic that earns them. This page explains what each layer is, when the pipeline moves to the next one, and how the layers combine into a single verdict.
The three layers¶
L1 — native detectors¶
Rule-based Rust detectors that need no model assets and are available fully offline. Runtime depends on input size and enabled rules; large inputs take milliseconds. L1 covers deterministic and pattern-based risks: prompt-injection phrasings, obfuscation tricks, secrets and destructive-operation patterns (DLP), format-validated PII (email, IBAN, credit card, …), and MCP tool-policy checks.
L1 runs for categories that implement native detectors when its execution gate is enabled. See Native detectors for the full catalogue.
L2 — NTDB model packages¶
Lightweight text classifiers in the Patronus NTDB format: the shared mmBERT tokenizer and
static embedder plus small ONNX heads and aggregators, packaged with a manifest.json. The
official L2 packages share this representation with compatible L3 models as well as with one
another, so adding categories is cheap and promotion does not start from raw text again. L2
answers in milliseconds and refines the L1 verdict with a learned classifier.
L2 promotion is decided by the NTDB package's promote-router. Separately, classifier verdicts
use a configurable final-decision threshold profile — see
ntdb_operating_point.
L3 — full ONNX transformers¶
Full transformer models (ModernBERT / mmBERT family), exported to ONNX and quantized. They are the most accurate and the most expensive. Which L3 models a gateway holds is fixed by configuration; they are kept resident in RAM (idle-TTL evicted) and executed by a background worker, not on the request path. L3 answers in tens of milliseconds and makes the final call for cases L2 could not resolve confidently.
L3 consumes the exact token IDs of the promoted L2 chunks. Both layers use the same 254-content-token chunks, with BOS/EOS and padding added for the 256-position L3 input. Invalid chunks report an error; there is no tokenization fallback.
Escalation: how a scan moves up¶
flowchart LR
IN([text]) --> L1
L1 --> L2{L2 assets<br/>ready?}
L2 -- no --> OUT1([L1 result])
L2 -- yes --> L2R[L2 classifier]
L2R --> P{promote?}
P -- no --> OUT2([L2 result])
P -- yes --> FB[[publish L2 fallback]]
FB --> W[L3 worker]
W --> OUT3([final L3 result])
- L1 always runs. Its verdict is the guaranteed baseline.
- L2 runs when its assets are cached and the level allows it (
max_level≥ L2). The L2 classifier produces a refined verdict. - Promotion to L3 happens when L2 is not confident enough to settle the case on its own
and the level allows it (
max_level= L3). The exact promotion behavior depends on the operating point. - On promotion, the pipeline immediately publishes the L2 fallback result, then enqueues an L3 job. The final L3 result is published later by the worker.
Because the L2 fallback is published first, a caller always has a usable answer quickly, even while the transformer is still running — and if L3 fails or times out, that fallback is what remains.
max_level is a hard ceiling¶
max_level ("l1", "l2", "l3") is the highest layer the gateway may ever use. It caps
escalation regardless of anything else:
max_level |
L1 | L2 | L3 |
|---|---|---|---|
l1 |
✅ | — | — |
l2 |
✅ | ✅ (if cached) | — |
l3 |
✅ | ✅ (if cached) | ✅ (if cached, on promotion) |
Execution gates can further disable individual levels or specific detectors below the ceiling, per request, without changing it.
Special pipelines¶
- L1-only categories (e.g.
pii) never load models and never promote. - L3-only pipelines (
dynamic-pii) enqueue directly to the worker and have no L1/L2 fallback. They can emit a non-authoritative first-entity preview before their completed result.
See Categories for each category's layer support.
Long text and windowing¶
Text is split into disjoint UTF-8 windows of at most 128 KiB, each tokenized once by compact mmBERT. L2 forms chunks of at most 254 content tokens; L3 reuses them unchanged. Tokenizer input memory is bounded per window; retained chunks and results still grow with document length. With representative clustering enabled (off by default) it groups near-duplicate windows by similarity, runs only cluster representatives, and propagates their verdict to the rest, so most windows never reach the model; early exit stops a head once its aggregate can no longer change. Aggregation strategy, clustering, and early exit are tunable per pipeline — see the L3 worker policy. Window and chunk counts are surfaced in the local benchmark's load report.
Degradation contract¶
Every failure path degrades to the best available result rather than throwing:
- L2 assets missing → return the L1 result.
- L3 asset missing / not promoted → return the L2 result.
- L3 inference error or timeout → return the published L2 fallback.
Failures are reported as structured SecurityFailure entries (stage + kind) on the terminal
event of the async API, so you can observe degradation without it breaking the scan. See the
Result schema.