Result schema¶
What the async queue (enqueue + consume_next_event) and the blocking scan_* helpers
return. Python names are shown; the Rust types (SecurityScanResult, LayerResult,
EvaluationResult, QueuedSecurityEvent, SecurityFailure) are in the
Rust API reference.
Scan result¶
A single result — the result field of a queue result event, or one element of the list the
blocking scan_* helpers return — is a dictionary describing one category's verdict:
[
{
"category": "dlp",
"class_name": "safe",
"confidence": 1.0,
"level": "L1",
"model": "native:dlp",
"duration_ms": 0.4,
"evidence_spans": [],
"label_scores": [],
"layers": [
{
"level": "L1",
"layer_type": "native",
"class_name": "safe",
"confidence": 1.0,
"matched": True,
"duration_ms": 0.4,
"thresholds": {},
"details": {},
}
],
}
]
| Field | Type | Meaning |
|---|---|---|
category |
str | The scan category this result belongs to. |
class_name |
str | Winning class (a category-specific label, or safe). |
confidence |
float | Confidence in class_name, 0.0–1.0. |
level |
str | The level that produced the winning verdict: L1, L2, or L3. |
model |
str | The producing scanner, e.g. native:dlp, external:<id>, or a model id. |
duration_ms |
float | Wall-clock time spent producing this result. |
decision |
dict | Structured classifier decision envelope for terminal model-backed classifier results. Python omits this key when no authoritative classifier decision exists; Rust exposes SecurityScanResult.decision: Option<DecisionEnvelope>. |
evidence_spans |
list | Exact matched spans from span-producing native scanners and dynamic PII; empty for safe/model-only results. |
label_scores |
list | Per-label scores for multi-label heads (e.g. tool_tags); each entry is {label, confidence, matched}. Empty for single-label results. |
layers |
list | Per-layer breakdown of everything that ran for this category. |
Layer entry¶
Each element of layers records one layer's output:
| Field | Type | Meaning |
|---|---|---|
level |
str | L1 / L2 / L3. |
layer_type |
str | e.g. native, or the model layer type. |
class_name |
str | This layer's class. |
confidence |
float | This layer's confidence. |
matched |
bool | Whether this layer produced a positive match. |
duration_ms |
float | Wall-clock time spent in this layer. |
thresholds |
dict | Thresholds applied at this layer (operating point, etc.). |
details |
dict | Layer-specific extra detail. Native PII/DLP layers can expose non-blocking l1_anchors; NTDB L2 layers add L3 promotion context; L3 layers add chunk execution metadata. The stable policy contract is decision, not free-form details. |
Decision envelope¶
Model-backed classifier pipelines (injection, threat, routing, sensitive_document,
tool_class, tool_action, and tool_tags) apply the configured final-decision threshold profile
after raw L2/L3 scoring. The top-level class_name / confidence is the accepted final verdict.
When no candidate passes its threshold, the result falls back to the pipeline default class
(benign, unknown, other, etc.) and uses the model's default-class confidence when available,
or 0.0 otherwise.
Terminal classifier results expose that policy input and Ark's calibrated recommendation under
decision. decision.is_some() is the lifecycle marker for an authoritative classifier policy
input. Early L2 results that still contain an l3_pending layer, provisional events, progress
events, and result-preview events leave decision unset (None in Rust) even when they carry
classifier-looking scores.
{
"schema_version": "ark.decision.v1",
"final_result": {"class_name": "benign", "confidence": 0.0, "source": "default"},
"decision_candidate": {
"source": "l2",
"class_name": "instruction_override",
"confidence": 0.5640919208526611,
"acceptance_threshold": 0.86471,
"accepted": False,
"evidence": None,
},
"recommendation": {
"accepted": False,
"final_arbitration": "default",
"operating_point": "best_f1",
"acceptance_threshold": 0.86471,
},
"candidates": [
{
"source": "l2",
"class_name": "instruction_override",
"confidence": 0.5640919208526611,
"acceptance_threshold": 0.86471,
"accepted": False,
"evidence": None,
}
],
"terminality": {"completion": "complete", "degraded": False, "degradation_reason": None},
"provenance": {
"ark_version": "0.1.7",
"schema_version": "ark.decision.v1",
"model": "unified-v3-threat",
},
}
| Field | Type | Meaning |
|---|---|---|
schema_version |
str | Decision-envelope schema version. Currently ark.decision.v1. |
final_result |
dict | The final Ark verdict after threshold arbitration: {class_name, confidence, source}. |
decision_candidate |
dict | null | Canonical policy input. It is the selected accepted or rejected candidate from L1/L2/L3/Union arbitration. None when no valid candidate exists. |
recommendation |
dict | Ark's calibrated default recommendation. accepted is false only for final_arbitration: "default". |
candidates |
list | All typed L1, L2, L3, and Union candidates available to arbitration. |
terminality |
dict | Completion state for the result: completion, degraded, and optional degradation_reason. |
provenance |
dict | Minimal source provenance: ark_version, schema_version, and model. |
Candidate entries have this shape:
| Field | Type | Meaning |
|---|---|---|
source |
str | l1, l2, l3, or union. final_result.source may also be default. |
class_name |
str | Candidate class before downstream Patronus policy. |
confidence |
float | Candidate confidence from that source. |
acceptance_threshold |
float | Ark's calibrated acceptance threshold for this candidate's source/class/operating point. |
accepted |
bool | Whether this candidate passed Ark's calibrated acceptance threshold. |
evidence |
dict | null | Extra candidate evidence. Union candidates include l2_weight, l3_weight, l2_confidence, and l3_confidence. |
Example: if Threat L2 predicts instruction_override at 0.5640919, but the selected threshold
profile requires a higher L2 confidence, the top-level result is the default benign class while
decision.decision_candidate retains the rejected l2 candidate with
acceptance_threshold: 0.86471 and decision.recommendation.accepted: False.
label_scores[].matched is layer-local model output metadata. It is not threshold acceptance and
must not be treated as equivalent to decision.recommendation.accepted or
decision.candidates[].accepted.
Candidate confidences are calibrated for Ark's bundled threshold profiles. Treat them in the
context of their category, source, model, and operating_point; do not compare scores across
unrelated sources or model versions without their matching thresholds.
Native L1 components¶
Built-in L1 matchers emit components directly from regex captures, lexical/structural
relationships, or decoded payloads. PII and DLP expose them under
layers[].details.matched_rules[].components, alongside the rule ID and finding range.
Components include component_id, explanation, start_byte, end_byte, and
span_precision. Contextual identifiers retain an anchor prefix/suffix and the validated
value. Injection's candidate features retain a complete rule_match plus its anchor
components; structural producer features retain their structural kind. Anchor decomposition
does not multiply the completed rule's scoring weight.
exact refers to original-text offsets, including mapped Unicode normalization.
transformed_source identifies the source container of a decoded payload when a narrower
character mapping is unavailable; the explanation identifies the decoded match.
Evidence spans¶
Native PII, all built-in DLP producers, accepted native Injection findings, and
dynamic-pii entities populate evidence_spans with original-text offsets:
for span in result["evidence_spans"]:
print(span["label"], span["text"], span["score"], span["start_byte"], span["end_byte"])
Each span carries the matched label, the matched text, a score, and both byte and
character offsets (start_byte/end_byte/start_char/end_char). Safe native results
leave evidence_spans empty.
PII spans with different labels may overlap: each matching class is retained, including a
numeric IBAN substring that also passes the credit-card validator. The primary class_name
does not enumerate every match; consume evidence_spans for all detected classes. Overlapping
matches within the same PII label are deduplicated.
Native PII and DLP scans do not compute or return the separate diagnostic context anchors by default.
Enable diagnostic context explicitly with Rust ScanGateMatrix.explain = true, Python
execution_gates={"explain": True}, or worker API gates: {explain: true}. Explained layers
can expose context under details.l1_anchors; these anchors are not findings:
{
"kind": "anchor",
"anchor_kind": "lexical",
"category": "date_of_birth",
"strength": "strong",
"text": "Geburtsdatum",
"start_byte": 18,
"end_byte": 30,
"start_char": 18,
"end_char": 30
}
Consumers must base immediate findings on evidence_spans and the result decision, not on an
anchor alone. Anchor metadata is diagnostic only and does not affect detection. Each native
layer includes at most 12 anchors within a 4-KiB serialized metadata budget; omitted anchors
set details.l1_anchors_truncated to true. Match text is previewed at most 96 UTF-8 bytes;
text_truncated: true marks a shortened preview, while all offsets still cover the full match.
Dynamic PII L3 layers expose details.inference_groups. Each entry identifies a stable base or
conditional GLiNER call, its labels, the conditional rule index, and optional byte ranges inherited
from matching final source-pipeline chunks. This is diagnostic provenance; the public findings
remain the merged evidence_spans, with the highest-scoring exact duplicate retained.
For registered native injection findings, the span label is the stable Ark rule ID. The
corresponding layer details contain an ordered matched_rules list with the Ark ID, optional
upstream ID, family, severity, description, source revision, byte offsets, and span_precision.
Data-driven regex rules use exact; procedural detectors currently use a localized clause or
bounded window, and a relationship assembled from independently matched components uses
composed. A provenance weight, when present, is metadata and not an Ark decision threshold.
Source-derived rules also expose references for secondary pinned sources, while source,
source_revision, source_license, upstream_id, and adaptation identify the primary origin
and Ark-specific narrowing.
The aggregated native Injection result exposes layers[].details.l1_candidates for accepted and
rejected candidates. Each candidate has a deterministic ID derived from its original-document
byte span, byte and character offsets, contributing producers, rule IDs and families, maximum
severity, calibrated score, threshold, acceptance result, and typed features:
{
"candidate_id": "injection:l1:18:57",
"category": "injection",
"start_byte": 18,
"end_byte": 57,
"start_char": 18,
"end_char": 57,
"rule_ids": ["ark.injection.override.hierarchy"],
"rule_severities": {"ark.injection.override.hierarchy": "critical"},
"families": ["instruction_override"],
"max_severity": "critical",
"producers": ["native:instruction_override"],
"score": 0.91,
"acceptance_threshold": 0.85,
"accepted": true,
"score_version": "injection-l1-0.1.6",
"features": [
{
"feature_id": "rule:ark.injection.override.hierarchy:18:57",
"kind": "rule_match",
"value": 1.0,
"explanation": "Invalidates or replaces a prior instruction hierarchy",
"start_byte": 18,
"end_byte": 57,
"span_precision": "clause",
"provenance": {
"rule_id": "ark.injection.override.hierarchy",
"source": "ark-native",
"source_revision": "71ff48e513ffee7810b29704e4cd9d4715aeaebd"
}
}
]
}
The separately gateable internal native:injection_structural producer uses the same candidate
contract. It may create a candidate without a flat catalog match.
Its relationship ID remains in rule_ids, while features[].kind is
structural; each feature has the exact span of one required component, such
as a context override, instruction-hierarchy reference, disclosure action, or
sensitive instruction object. The candidate span is the smallest
original-document region containing all required components.
The native:injection_l1 aggregate scores each merged candidate. Accepted candidates create
public finding spans and use source: "l1"; rejected candidates remain visible in
decision.candidates with accepted: false while the top-level result stays safe. Individual
native Injection producer verdicts are no longer returned as separate public results. Registered
external L1 detectors remain separate and unchanged.
Async queue events¶
consume_next_event(timeout) returns one event dict at a time (or None on timeout).
consume_events(timeout) yields them. There are up to four event_types: result and
finished always, plus progress and provisional when L3 progress reporting is enabled
(execution_gates.l3.progress, off by default). The blocking scan_* helpers only ever return
result / finished.
result¶
One request can emit several result events. L1 results are visible as soon as L1
finishes; L2 and a later promoted L3 result follow independently.
Promoting NTDB L2 layers expose details.l3_candidates (each entry: byte span,
promote_score, promote_threshold, source_pipeline, source_model, l2_class) and
details.l2_chunk_outputs (the per-chunk L2 outputs retained for aggregation, with the raw
embeddings stripped). Under l3_strategy="multi", candidates from every promoted pipeline in the
coalesced run are deduplicated by exact span match — when two heads promote the same span,
source_pipeline / source_model / l2_class become comma-joined unions; under the default
l3_strategy="dedicated" each pipeline's candidates stay separate and this merge never runs. The
L3 worker maps candidate spans onto tokenizer-bounded windows; with representative
clustering enabled it groups near-duplicate windows, infers only cluster representatives, and
propagates their verdict to the rest (see
L3 worker policy). An empty or unusable candidate list
deliberately falls back to full-text L3.
For Package-v4 Union decisions, the final decision_candidate may contain chunk_evidence with
the aggregation method and the contributing document chunks. This provenance never changes the
scan verdict.
progress¶
Opt-in, non-authoritative status while a long L3 scan resolves — only under the dedicated L3
strategy (l3_strategy="multi" never emits it):
{
"event_type": "progress",
"request_id": "…",
"progress": {
"category": "injection", "model": "…", "stage": "l3_chunk",
"completed_chunks": 3, "total_chunks": 8,
"inferred_chunks": 2, "propagated_chunks": 1, "cache_hits": 0,
"early_exit": False, "coverage": 0.375, "details": { … },
},
}
stage is one of l3_started, l3_chunk, l3_cluster_propagated, l3_early_exit. A progress
event carries no verdict — never treat it as a result.
provisional¶
Same shape as a result event's result, but an interim preview re-aggregated from the
chunks resolved so far. It is non-authoritative and may be superseded by later result /
finished events; only emitted when execution_gates.l3.progress is set to "provisional".
finished¶
Exactly one terminal event follows all results for a request:
{
"event_type": "finished",
"request_id": "…",
"completion": "…", # "complete" | "degraded" | "failed"
"failures": [ … ], # structured failures, if any
}
completion is one of complete (all planned work succeeded), degraded (some layer failed
but a lower-layer result was delivered), or failed (no usable result). Consuming finished
removes all library state for that request ID. Correlate every event by request_id.
Failures¶
Each failures entry is a structured SecurityFailure dict:
| Field | Type | Meaning |
|---|---|---|
stage |
str | warmup, asset, scanner, inference, queue, or worker. |
kind |
str | not_ready, missing_asset, integrity_failure, initialization_failure, inference_failure, timeout, queue_full, worker_unavailable, or internal. |
level |
str | null | The level that failed (L1/L2/L3), if applicable. |
detector_id |
str | null | The specific detector or model that failed, if applicable. |
retryable |
bool | Whether the failure is transient and could succeed on retry. |
message |
str | Human-readable description. |
A failure does not throw during scanning — the scan
degrades to the best available
lower-layer result and reports the failure here. (warmup() itself is the exception: a missing
required asset there raises rather than degrades.)
Request introspection¶
| Method | Returns |
|---|---|
has_request(request_id) |
Whether the gateway still tracks this request. |
request_state(request_id) |
Current state dict, or None. |
is_finished(request_id) |
True / False, or None if unknown. |
runtime_readiness() |
Initialized runtime state (levels ready, per stage/kind). |
Request state reflects SecurityRequestState — whether any planned scanner or promoted L3 job
can still publish an event.