Skip to content

Result schema

What the async queue (enqueue + consume_next_event) and the blocking scan_* helpers return. Python names are shown; the Rust types (SecurityScanResult, LayerResult, EvaluationResult, QueuedSecurityEvent, SecurityFailure) are in the Rust API reference.

Scan result

A single result — the result field of a queue result event, or one element of the list the blocking scan_* helpers return — is a dictionary describing one category's verdict:

[
    {
        "category": "dlp",
        "class_name": "safe",
        "confidence": 1.0,
        "level": "L1",
        "model": "native:dlp",
        "duration_ms": 0.4,
        "evidence_spans": [],
        "label_scores": [],
        "layers": [
            {
                "level": "L1",
                "layer_type": "native",
                "class_name": "safe",
                "confidence": 1.0,
                "matched": True,
                "duration_ms": 0.4,
                "thresholds": {},
                "details": {},
            }
        ],
    }
]
Field Type Meaning
category str The scan category this result belongs to.
class_name str Winning class (a category-specific label, or safe).
confidence float Confidence in class_name, 0.0–1.0.
level str The level that produced the winning verdict: L1, L2, or L3.
model str The producing scanner, e.g. native:dlp, external:<id>, or a model id.
duration_ms float Wall-clock time spent producing this result.
decision dict Structured classifier decision envelope for terminal model-backed classifier results. Python omits this key when no authoritative classifier decision exists; Rust exposes SecurityScanResult.decision: Option<DecisionEnvelope>.
evidence_spans list Exact matched spans from span-producing native scanners and dynamic PII; empty for safe/model-only results.
label_scores list Per-label scores for multi-label heads (e.g. tool_tags); each entry is {label, confidence, matched}. Empty for single-label results.
layers list Per-layer breakdown of everything that ran for this category.

Layer entry

Each element of layers records one layer's output:

Field Type Meaning
level str L1 / L2 / L3.
layer_type str e.g. native, or the model layer type.
class_name str This layer's class.
confidence float This layer's confidence.
matched bool Whether this layer produced a positive match.
duration_ms float Wall-clock time spent in this layer.
thresholds dict Thresholds applied at this layer (operating point, etc.).
details dict Layer-specific extra detail. Native PII/DLP layers can expose non-blocking l1_anchors; NTDB L2 layers add L3 promotion context; L3 layers add chunk execution metadata. The stable policy contract is decision, not free-form details.

Decision envelope

Model-backed classifier pipelines (injection, threat, routing, sensitive_document, tool_class, tool_action, and tool_tags) apply the configured final-decision threshold profile after raw L2/L3 scoring. The top-level class_name / confidence is the accepted final verdict. When no candidate passes its threshold, the result falls back to the pipeline default class (benign, unknown, other, etc.) and uses the model's default-class confidence when available, or 0.0 otherwise.

Terminal classifier results expose that policy input and Ark's calibrated recommendation under decision. decision.is_some() is the lifecycle marker for an authoritative classifier policy input. Early L2 results that still contain an l3_pending layer, provisional events, progress events, and result-preview events leave decision unset (None in Rust) even when they carry classifier-looking scores.

{
    "schema_version": "ark.decision.v1",
    "final_result": {"class_name": "benign", "confidence": 0.0, "source": "default"},
    "decision_candidate": {
        "source": "l2",
        "class_name": "instruction_override",
        "confidence": 0.5640919208526611,
        "acceptance_threshold": 0.86471,
        "accepted": False,
        "evidence": None,
    },
    "recommendation": {
        "accepted": False,
        "final_arbitration": "default",
        "operating_point": "best_f1",
        "acceptance_threshold": 0.86471,
    },
    "candidates": [
        {
            "source": "l2",
            "class_name": "instruction_override",
            "confidence": 0.5640919208526611,
            "acceptance_threshold": 0.86471,
            "accepted": False,
            "evidence": None,
        }
    ],
    "terminality": {"completion": "complete", "degraded": False, "degradation_reason": None},
    "provenance": {
        "ark_version": "0.1.7",
        "schema_version": "ark.decision.v1",
        "model": "unified-v3-threat",
    },
}
Field Type Meaning
schema_version str Decision-envelope schema version. Currently ark.decision.v1.
final_result dict The final Ark verdict after threshold arbitration: {class_name, confidence, source}.
decision_candidate dict | null Canonical policy input. It is the selected accepted or rejected candidate from L1/L2/L3/Union arbitration. None when no valid candidate exists.
recommendation dict Ark's calibrated default recommendation. accepted is false only for final_arbitration: "default".
candidates list All typed L1, L2, L3, and Union candidates available to arbitration.
terminality dict Completion state for the result: completion, degraded, and optional degradation_reason.
provenance dict Minimal source provenance: ark_version, schema_version, and model.

Candidate entries have this shape:

Field Type Meaning
source str l1, l2, l3, or union. final_result.source may also be default.
class_name str Candidate class before downstream Patronus policy.
confidence float Candidate confidence from that source.
acceptance_threshold float Ark's calibrated acceptance threshold for this candidate's source/class/operating point.
accepted bool Whether this candidate passed Ark's calibrated acceptance threshold.
evidence dict | null Extra candidate evidence. Union candidates include l2_weight, l3_weight, l2_confidence, and l3_confidence.

Example: if Threat L2 predicts instruction_override at 0.5640919, but the selected threshold profile requires a higher L2 confidence, the top-level result is the default benign class while decision.decision_candidate retains the rejected l2 candidate with acceptance_threshold: 0.86471 and decision.recommendation.accepted: False.

label_scores[].matched is layer-local model output metadata. It is not threshold acceptance and must not be treated as equivalent to decision.recommendation.accepted or decision.candidates[].accepted.

Candidate confidences are calibrated for Ark's bundled threshold profiles. Treat them in the context of their category, source, model, and operating_point; do not compare scores across unrelated sources or model versions without their matching thresholds.

Native L1 components

Built-in L1 matchers emit components directly from regex captures, lexical/structural relationships, or decoded payloads. PII and DLP expose them under layers[].details.matched_rules[].components, alongside the rule ID and finding range. Components include component_id, explanation, start_byte, end_byte, and span_precision. Contextual identifiers retain an anchor prefix/suffix and the validated value. Injection's candidate features retain a complete rule_match plus its anchor components; structural producer features retain their structural kind. Anchor decomposition does not multiply the completed rule's scoring weight.

exact refers to original-text offsets, including mapped Unicode normalization. transformed_source identifies the source container of a decoded payload when a narrower character mapping is unavailable; the explanation identifies the decoded match.

Evidence spans

Native PII, all built-in DLP producers, accepted native Injection findings, and dynamic-pii entities populate evidence_spans with original-text offsets:

for span in result["evidence_spans"]:
    print(span["label"], span["text"], span["score"], span["start_byte"], span["end_byte"])

Each span carries the matched label, the matched text, a score, and both byte and character offsets (start_byte/end_byte/start_char/end_char). Safe native results leave evidence_spans empty.

PII spans with different labels may overlap: each matching class is retained, including a numeric IBAN substring that also passes the credit-card validator. The primary class_name does not enumerate every match; consume evidence_spans for all detected classes. Overlapping matches within the same PII label are deduplicated.

Native PII and DLP scans do not compute or return the separate diagnostic context anchors by default. Enable diagnostic context explicitly with Rust ScanGateMatrix.explain = true, Python execution_gates={"explain": True}, or worker API gates: {explain: true}. Explained layers can expose context under details.l1_anchors; these anchors are not findings:

{
  "kind": "anchor",
  "anchor_kind": "lexical",
  "category": "date_of_birth",
  "strength": "strong",
  "text": "Geburtsdatum",
  "start_byte": 18,
  "end_byte": 30,
  "start_char": 18,
  "end_char": 30
}

Consumers must base immediate findings on evidence_spans and the result decision, not on an anchor alone. Anchor metadata is diagnostic only and does not affect detection. Each native layer includes at most 12 anchors within a 4-KiB serialized metadata budget; omitted anchors set details.l1_anchors_truncated to true. Match text is previewed at most 96 UTF-8 bytes; text_truncated: true marks a shortened preview, while all offsets still cover the full match.

Dynamic PII L3 layers expose details.inference_groups. Each entry identifies a stable base or conditional GLiNER call, its labels, the conditional rule index, and optional byte ranges inherited from matching final source-pipeline chunks. This is diagnostic provenance; the public findings remain the merged evidence_spans, with the highest-scoring exact duplicate retained.

For registered native injection findings, the span label is the stable Ark rule ID. The corresponding layer details contain an ordered matched_rules list with the Ark ID, optional upstream ID, family, severity, description, source revision, byte offsets, and span_precision. Data-driven regex rules use exact; procedural detectors currently use a localized clause or bounded window, and a relationship assembled from independently matched components uses composed. A provenance weight, when present, is metadata and not an Ark decision threshold. Source-derived rules also expose references for secondary pinned sources, while source, source_revision, source_license, upstream_id, and adaptation identify the primary origin and Ark-specific narrowing.

The aggregated native Injection result exposes layers[].details.l1_candidates for accepted and rejected candidates. Each candidate has a deterministic ID derived from its original-document byte span, byte and character offsets, contributing producers, rule IDs and families, maximum severity, calibrated score, threshold, acceptance result, and typed features:

{
  "candidate_id": "injection:l1:18:57",
  "category": "injection",
  "start_byte": 18,
  "end_byte": 57,
  "start_char": 18,
  "end_char": 57,
  "rule_ids": ["ark.injection.override.hierarchy"],
  "rule_severities": {"ark.injection.override.hierarchy": "critical"},
  "families": ["instruction_override"],
  "max_severity": "critical",
  "producers": ["native:instruction_override"],
  "score": 0.91,
  "acceptance_threshold": 0.85,
  "accepted": true,
  "score_version": "injection-l1-0.1.6",
  "features": [
    {
      "feature_id": "rule:ark.injection.override.hierarchy:18:57",
      "kind": "rule_match",
      "value": 1.0,
      "explanation": "Invalidates or replaces a prior instruction hierarchy",
      "start_byte": 18,
      "end_byte": 57,
      "span_precision": "clause",
      "provenance": {
        "rule_id": "ark.injection.override.hierarchy",
        "source": "ark-native",
        "source_revision": "71ff48e513ffee7810b29704e4cd9d4715aeaebd"
      }
    }
  ]
}

The separately gateable internal native:injection_structural producer uses the same candidate contract. It may create a candidate without a flat catalog match. Its relationship ID remains in rule_ids, while features[].kind is structural; each feature has the exact span of one required component, such as a context override, instruction-hierarchy reference, disclosure action, or sensitive instruction object. The candidate span is the smallest original-document region containing all required components.

The native:injection_l1 aggregate scores each merged candidate. Accepted candidates create public finding spans and use source: "l1"; rejected candidates remain visible in decision.candidates with accepted: false while the top-level result stays safe. Individual native Injection producer verdicts are no longer returned as separate public results. Registered external L1 detectors remain separate and unchanged.

Async queue events

consume_next_event(timeout) returns one event dict at a time (or None on timeout). consume_events(timeout) yields them. There are up to four event_types: result and finished always, plus progress and provisional when L3 progress reporting is enabled (execution_gates.l3.progress, off by default). The blocking scan_* helpers only ever return result / finished.

result

{
    "event_type": "result",
    "request_id": "…",
    "result": { …a scan result dict, as above… },
}

One request can emit several result events. L1 results are visible as soon as L1 finishes; L2 and a later promoted L3 result follow independently.

Promoting NTDB L2 layers expose details.l3_candidates (each entry: byte span, promote_score, promote_threshold, source_pipeline, source_model, l2_class) and details.l2_chunk_outputs (the per-chunk L2 outputs retained for aggregation, with the raw embeddings stripped). Under l3_strategy="multi", candidates from every promoted pipeline in the coalesced run are deduplicated by exact span match — when two heads promote the same span, source_pipeline / source_model / l2_class become comma-joined unions; under the default l3_strategy="dedicated" each pipeline's candidates stay separate and this merge never runs. The L3 worker maps candidate spans onto tokenizer-bounded windows; with representative clustering enabled it groups near-duplicate windows, infers only cluster representatives, and propagates their verdict to the rest (see L3 worker policy). An empty or unusable candidate list deliberately falls back to full-text L3.

For Package-v4 Union decisions, the final decision_candidate may contain chunk_evidence with the aggregation method and the contributing document chunks. This provenance never changes the scan verdict.

progress

Opt-in, non-authoritative status while a long L3 scan resolves — only under the dedicated L3 strategy (l3_strategy="multi" never emits it):

{
    "event_type": "progress",
    "request_id": "…",
    "progress": {
        "category": "injection", "model": "…", "stage": "l3_chunk",
        "completed_chunks": 3, "total_chunks": 8,
        "inferred_chunks": 2, "propagated_chunks": 1, "cache_hits": 0,
        "early_exit": False, "coverage": 0.375, "details": { … },
    },
}

stage is one of l3_started, l3_chunk, l3_cluster_propagated, l3_early_exit. A progress event carries no verdict — never treat it as a result.

provisional

Same shape as a result event's result, but an interim preview re-aggregated from the chunks resolved so far. It is non-authoritative and may be superseded by later result / finished events; only emitted when execution_gates.l3.progress is set to "provisional".

finished

Exactly one terminal event follows all results for a request:

{
    "event_type": "finished",
    "request_id": "…",
    "completion": "…",     # "complete" | "degraded" | "failed"
    "failures": [ … ],     # structured failures, if any
}

completion is one of complete (all planned work succeeded), degraded (some layer failed but a lower-layer result was delivered), or failed (no usable result). Consuming finished removes all library state for that request ID. Correlate every event by request_id.

Failures

Each failures entry is a structured SecurityFailure dict:

Field Type Meaning
stage str warmup, asset, scanner, inference, queue, or worker.
kind str not_ready, missing_asset, integrity_failure, initialization_failure, inference_failure, timeout, queue_full, worker_unavailable, or internal.
level str | null The level that failed (L1/L2/L3), if applicable.
detector_id str | null The specific detector or model that failed, if applicable.
retryable bool Whether the failure is transient and could succeed on retry.
message str Human-readable description.

A failure does not throw during scanning — the scan degrades to the best available lower-layer result and reports the failure here. (warmup() itself is the exception: a missing required asset there raises rather than degrades.)

Request introspection

Method Returns
has_request(request_id) Whether the gateway still tracks this request.
request_state(request_id) Current state dict, or None.
is_finished(request_id) True / False, or None if unknown.
runtime_readiness() Initialized runtime state (levels ready, per stage/kind).

Request state reflects SecurityRequestState — whether any planned scanner or promoted L3 job can still publish an event.