Detector maturity and release gates (G1–G3)
- Status: Proposal. Risk, product, SRE, privacy and legal owners have not yet approved it.
- Date: 2026-10-03 (Australia/Sydney)
- First profiles: HPI-I (
validated_structured) and labelled Victorian Student Number (context_bound_structured) - Machine-readable draft:
test/gates/detector-gates.v1alpha1.json, validated bytest/gates/detector-gates.schema.json - Decisions: implements the evidence structure for OD-004 and OD-019. It does not settle either decision.
- Current state: the HPI-I (
builtin.pii.au.hpi-i) and VSN (builtin.pii.au-vic.vsn) detectors ship as experimental version0.1.0, disabled by default (G0). Theau-baselineobserveandredact-experimentalprofiles exist for synthetic traffic, but no G1 or G2 evidence has been produced yet. See the micro-slice record.
1. Purpose and boundary
This document sets the evidence a detector version needs before a policy may use it for a stronger action. It extends the conformance and benchmarking strategy and keeps the separation in ADR 0002: detectors report findings, and policy alone decides whether to record, redact or block.
A gate certifies one tuple:
(detector_id, detector_version, category, detector_class,
protocol paths, locales/jurisdictions, maximum policy action)Passing G3 means the detector is eligible for blocking policies. It does not mean any policy must block. Each policy owner still decides by route and stage, including OD-005, OD-006 and OD-010.
2. Gate ladder
| Level | Name | Highest policy action allowed | Catalogue lifecycle (ADR 0011) | Scope |
|---|---|---|---|---|
| G0 | Development | none (tests and dev snapshots only) | development | local, CI, synthetic data |
| G1 | Observe | record, required: false | experimental | any approved binding; no effect on traffic |
| G2 | Experimental redact | redact (and record) | experimental | explicit opt-in canary bindings with rollback |
| G3 | Production block | block, redact, record | stable | general bindings, subject to policy-owner choice |
Rules:
- Levels are cumulative. Every Gn criterion stays in force at Gn+1.
- A new detector version must be re-evaluated against its current level. ADR 0011 already classifies expanded recognition as a minor change that needs a decision diff review. A detector version that fails is not promoted. The previously promoted version remains active.
- Demotion: if an invariant fails in CI or production, disable the affected
action immediately. ADR 0011 allows
unsupportedfor unsafe behaviour. If a monitored statistical bound fails, the detector drops one level, starting a review within an SLA that requires approval. - Under the ADR 0011 compiler and runtime-snapshot plan, policy compilation should eventually reject a rule whose action exceeds the referenced detector version's promoted gate. That is future compiler work. This document does not change runtime code.
3. Value-status legend
Each number has a status. Anything not marked I or D is not yet an
organizational commitment.
| Mark | Meaning | Who may change it |
|---|---|---|
| I invariant | Follows from the detector contract or an ADR. Any failure blocks the gate. No sampling or tolerance applies. | Only an ADR or contract change |
| D derived | A mathematical consequence of a chosen target and confidence level, for example minimum n. Recompute it if the target changes. | Automatically follows its inputs |
| P provisional | Reversible engineering default with a recorded rationale. Replace it with measured or approved values. | Detector and platform owners |
| R requires approval | Proposed value only. A G2/G3 claim cannot pass until the named owner approves it. | Risk + product (OD-004), SRE, privacy, legal as named |
| M measured-then-frozen | Set from the measured baseline at the previous gate plus approved headroom | Owner named for the metric |
| unset | No defensible number exists yet. The criterion is report-only until owners set it. | Named owner |
No confidence-score thresholds are proposed. Section 10 explains why. No number in this document came from an uncited "industry standard".
4. Detector-class distinction
ADR 0010 defines two structured classes. Their failure modes differ, so the gates differ.
| Aspect | validated_structured (HPI-I) | context_bound_structured (labelled VSN) |
|---|---|---|
| Evidence for a finding | Authoritative grammar, plus checksum where published | Approved exact label or structured field, plus jurisdiction where relevant, plus candidate grammar |
| Main precision risk | Structural collisions: sibling prefixes (IHI/HPI-O), embedded digit runs, other Luhn numbers | Label ambiguity and binding errors. A bare value has no evidential value. |
| Main recall risk | Formatting outside the grammar, Unicode obfuscation | Label variants, field aliases, binding distance, tables and lists |
| Highest validation state | validated only when prefix + length + Luhn pass. This means structurally valid, not issued (no HI Service lookup; OD-016). | probable. With no checksum and no approved VSR lookup, validated is impossible. |
| Class invariant | Type exactness across the 80036x family; no finding on Luhn failure unless a label decision allows candidate (decision L-H2) | The same value without an approved label or field produces no finding |
| Statistical gates focus on | Benign digit-dense text, formatting variety, protocol paths | Label/binding variety, hard negatives with competing labels, observed precision |
| Score meaning (ADR 0010) | Local structural plausibility | Joint label, jurisdiction and candidate strength |
These are not interchangeable. A context-bound detector cannot get its precision
from checksum math, and a validated-structured detector cannot use label
evidence to exceed validated.
5. Corpus requirements
5.1 Partitions
Use the three partitions from the conformance strategy, plus purpose-specific suites:
| Suite | Purpose | Statistical? |
|---|---|---|
conformance | Specification cases authored from the cited source and label conventions. Every case must pass. | No. Invariant. |
development | Public synthetic cases for tuning and debugging | Reported only |
holdout | Private, fixed, independently authored cases, never used for tuning | Yes |
hard_negative | Near-miss families (section 11) | Yes, per family |
benign | Realistic documents without the target entity, deliberately digit-dense | Yes |
adversarial | Transformations of holdout base cases (section 7) | Invariant per disposition |
observed_traffic | Reviewed sample of observe-mode findings, only under an approved privacy process | Yes |
performance | Payload-size and concurrency matrix | Order-statistic / benchstat |
5.2 Unit of analysis: base cases, not variants
Deterministic generators can multiply one template into thousands of cases. These are not independent evidence. Miller's clustered-error analysis shows naive standard errors can understate uncertainty several-fold when items are related. Therefore:
- Statistical
ncounts independent base cases. A base case is an independently authored or independently sampled document with its own context. - Variants such as Unicode, separators, position and wrapper format are
invariant multipliers. Every variant of a passing base case must pass under
its disposition. Variants never increase
n. - For recall, the gating indicator is per base case: all gold target entities are found with exact type and span. Entity-level metrics are also reported. This keeps the binomial model valid without a cluster bootstrap. A cluster bootstrap over base cases may be reported alongside it.
5.3 Minimum sizes
Minimum sizes are derived from the target being demonstrated. Gate decisions use the one-sided 95% Clopper–Pearson bound. For zero failures, this is the exact form of the rule of three.
Minimum independent base cases n needed for the one-sided 95% upper bound on a
failure rate to fall at or below the target. Computed exactly. Applies to a miss
rate or FP rate.
| Target failure rate | 0 failures | ≤1 | ≤2 | ≤3 | ≤5 |
|---|---|---|---|---|---|
| 5% | 59 | 93 | 124 | 153 | 208 |
| 2% | 149 | 236 | 313 | 386 | 523 |
| 1% | 299 | 473 | 628 | 773 | 1,049 |
| 0.5% | 598 | 947 | 1,258 | 1,549 | 2,100 |
| 0.1% | 2,995 | 4,742 | 6,294 | 7,752 | 10,511 |
| 0.01% | 29,956 | 47,437 | 62,956 | 77,535 | 105,128 |
At 99% confidence, zero failures require 459 cases for 1% and 4,603 for 0.1%.
Proposed corpus minimums by gate. The counts are D given the targets in section 6. The targets are R.
| Suite | G1 | G2 | G3 |
|---|---|---|---|
| Conformance | 100% of the specification matrix (I) | same | same |
| Holdout positive base cases | ≥300 | ≥500 (allows ≤1 miss at 1%) | ≥1,000 (allows ≤1 miss at 0.5%) |
| Each critical slice (section 11) | ≥59 | ≥59 | ≥149 |
| Hard-negative base cases | ≥59 per family; ≥300 total | ≥299 per family | ≥299 per family and ≥2,995 total |
| Benign documents | ≥1,000 (P) | ≥2,995 | ≥2,995 synthetic plus observed-traffic evidence (section 6.5) |
| Bare-value contexts (context-bound only) | ≥1,000 distinct contexts × generated values (P); coverage, not statistics | same | same |
| Annotation | Two independent annotators per holdout/hard-negative case, adjudicated, versioned guideline | same | same, and an adjudication log is in the evidence bundle |
Report inter-annotator agreement (span-level strict F1 between annotators and Cohen's κ on case labels). Every disagreement must be adjudicated before the holdout is frozen. No κ threshold is proposed. For deterministic synthetic identifiers, disagreement is a label-guideline defect to fix, not noise to tolerate.
All corpora stay synthetic by default. Values must be generated. Never use real HPI-Is or VSNs. Any traffic-derived set follows the privacy and corpus register and OD-008/OD-009.
6. Quantitative criteria
Notation: LB and UB are one-sided 95% Clopper–Pearson lower and upper bounds.
"Strict" means exact category and exact code-point span, as in SemEval-2013
Task 9 ("strict" schema) and the
entity-based i2b2 2014 de-identification evaluation
(Stubbs et al., 2015). Report
exact-boundary-any-type, partial-overlap and character-level scores too. Gates
use only strict scores.
6.1 Exact type/span correctness
| ID | Criterion | G1 | G2 | G3 |
|---|---|---|---|---|
| C1 | Conformance strict pass rate | 100% (I) | 100% (I) | 100% (I) |
| C2 | Type confusion within the structural family (IHI/HPI-I/HPI-O; VSN vs TFN, ACN, CRN, VCAA Student Number) on conformance | 0 (I) | 0 (I) | 0 (I) |
| C3 | Holdout strict recall, base-case indicator | report, n≥300 | LB ≥ 0.99 (R) | LB ≥ 0.995 (R) |
| C4 | Holdout strict precision, per predicted-finding document | report | LB ≥ 0.99 (R) | LB ≥ 0.995 (R) |
| C5 | Strict recall in each critical slice | report, n≥59 | LB ≥ 0.95 (R) | LB ≥ 0.98 (R) |
| C6 | Evidence completeness: every finding carries the class's required evidence kinds and never contains the matched value | 100% (I) | 100% (I) | 100% (I) |
| C7 | Validation-state ceiling (section 4) | 0 violations (I) | 0 (I) | 0 (I) |
| C8 | Context-bound only: findings on bare-value suite | 0 (I) | 0 (I) | 0 (I) |
The machine-readable templates split C7 into class-specific checks: V1 means
validated without checksum evidence, and C9 means validated on a
context-bound finding.
The G2/G3 floors are deliberately close to 1. For a deterministic structured
detector on synthetic in-grammar data, an in-specification miss is a bug, not a
statistical tolerance. That case is C1. The holdout floors bound the residual
risk from realistic context, formatting and label variation that conformance
authors did not foresee. Risk owners may choose different floors per action.
Section 5.3 then gives the new n.
6.2 Hard negatives and benign pass
| ID | Criterion | G1 | G2 | G3 |
|---|---|---|---|---|
| H1 | Structural near-miss families with deterministic expected outcomes (wrong Luhn, sibling prefix, wrong length, embedded runs) | 0 FP (I) | 0 FP (I) | 0 FP (I) |
| H2 | FP rate per semantic hard-negative family (competing label, binding failure, ambiguous acronym) | report, ≥59/family | UB ≤ 0.01 per family (R) | UB ≤ 0.01 per family and UB ≤ 0.001 aggregate (R) |
| B1 | Benign pass rate: benign documents with no finding of the category | report | LB ≥ 0.999 (R) | LB ≥ 0.999 synthetic and observed-traffic precision O1 (R) |
Benign pass sits beside every recall result, as the conformance strategy requires. Synthetic precision depends on synthetic prevalence. Production precision needs observed-traffic review (section 6.5).
6.3 Redaction leakage and over-redaction
These are measured end to end. Detector findings pass through the transformer and adapter, including Portkey full-object reconstruction. An independent oracle then re-scans the output. The oracle is a separate reference implementation in the Python evaluator with normalization (NFKC, removal of default-ignorable code points, digit folding).
| ID | Criterion | G1 | G2 | G3 |
|---|---|---|---|---|
| R1 | Span completeness given detection: no gold value code point survives for any detected entity | report (simulated transform) | 0 residual (I) | 0 residual (I) |
| R2 | Full leakage: gold entity recoverable by the oracle after redaction | report | UB ≤ 0.01 (R) | UB ≤ 0.005 (R) |
| R3 | Any leakage: ≥1 gold value code point survives, excluding separators | report | UB ≤ 0.01 (R) | UB ≤ 0.005 (R) |
| R4 | Boundary over-extension: redaction outside gold span on conformance | 0 (I) | 0 (I) | 0 (I) |
| R5 | Over-redaction character rate on holdout positives (non-gold code points redacted ÷ non-gold code points) | report | report (ceiling unset; product) | ceiling required (R, unset) |
| R6 | Benign modification rate: benign documents with any modification attributed to the category | report | UB ≤ 0.001 (R) | UB ≤ 0.001 synthetic and observed (R) |
| R7 | Replacement objects schema-valid and idempotent; unrelated payload fields preserved | 100% (I) | 100% (I) | 100% (I) |
Leakage is defined against OD-010 semantics. If owners later allow partial masking (for example, keep the last three digits), R3 must exclude the permitted residual. R1 must then also use a documented masking rule.
6.4 Unicode and adversarial robustness
Each detector profile assigns every transformation in the catalogue one disposition:
must_detect: same strict result as the clean base case, with spans into the original text;must_not_detect: must produce no finding;declared_gap: published in detector documentation and capability metadata as a known coverage gap; never silently treated as clean;requires_decision: blocks G2 until an owner chooses a disposition.
| ID | Criterion | G1 | G2 | G3 |
|---|---|---|---|---|
| U1 | Span integrity on every transformed case: valid code-point offsets, no panic, no coverage loss, no timeout | 100% (I) | 100% (I) | 100% (I) |
| U2 | Transform-induced misses on must_detect | report | 0 (I given disposition) | 0 (I) |
| U3 | Findings on must_not_detect | report | 0 (I) | 0 (I) |
| U4 | Every catalogue entry has a disposition; no requires_decision remains | listed | complete (I) | complete (I) |
| U5 | Structured adversarial review round: humans or tools create bypass attempts beyond the catalogue | — | — | no open high-severity bypass without signed risk acceptance (R) |
| U6 | Fuzzing of detector and span application, with no crash or invariant violation | nightly budget (P) | nightly budget (P) | cumulative budget recorded (P) |
Base catalogue, extensible by profile. References: UAX #15, UTS #39, UTR #36:
| ID | Transformation |
|---|---|
| T01 | Emoji, CJK or combining marks before, after or between findings (offset asymmetry) |
| T02 | NFC/NFD/NFKC/NFKD variants of surrounding text and labels |
| T03 | Non-ASCII decimal digits (Nd: fullwidth U+FF10–FF19, Arabic-Indic, Devanagari, etc.) |
| T04 | Default-ignorable characters inside the value (U+200B/C/D, U+2060, U+FEFF, U+00AD) |
| T05 | Space-like separators inside the value (U+0020, U+00A0, U+2009, U+202F) |
| T06 | Dash-like separators (U+002D, U+2010–U+2015, U+2212) |
| T07 | Bidirectional controls (U+202A–U+202E, U+2066–U+2069) around or inside the value |
| T08 | Keycap or enclosed digits (digit + U+FE0F + U+20E3; circled digits) |
| T09 | Homoglyph or mixed-script labels (Cyrillic or Greek look-alikes in VSN, HPI-I) |
| T10 | Label case and punctuation variants (vsn, V.S.N., HPI I, HPII) |
| T11 | Line breaks, Markdown/table/code-fence/CSV wrappers, JSON string escapes (\u0038) |
| T12 | Value split across tokens, segments or messages |
| T13 | HTML entities, percent-encoding, Base64 or other encodings |
| T14 | Paraphrased or translated labels (for example, "student number issued in Victoria") |
Grapheme-boundary rule (proposed, decision U-1): a redaction span must end on an extended grapheme cluster boundary (UAX #29). This prevents redacting a digit while leaving an attached combining keycap (T08), and the reverse. Spans remain code-point offsets per ADR 0003. The rule only constrains where they may end.
6.5 Observed-traffic evidence (between gates)
Synthetic corpora cannot estimate production precision or prevalence. The transitions G1→G2 and G2→G3 need reviewed traffic evidence. If privacy owners do not approve review (OD-009/OD-008), promotion must instead record a signed risk acceptance stating that only synthetic evidence exists.
| ID | Criterion | G1→G2 | G2→G3 |
|---|---|---|---|
| O1 | Observed precision: uniformly random sample of observe/canary findings, reviewed by approved reviewers in a controlled environment, using keyed fingerprints where possible | LB ≥ 0.99, ≥299 reviewed or all findings (R) | LB ≥ 0.995, ≥598 reviewed (R) |
| O2 | Observation window | ≥14 days covering two weekly cycles (P) | ≥14 days of canary (P) |
| O3 | Stratified review of near misses: label-adjacent unbound values (context-bound) and in-grammar-but-failed values (validated) | report | report; any confirmed missed true positive becomes a holdout case |
| O4 | Coverage health: timed_out/failed/partial rate for this check | report | ≤ SRE budget (R, unset) |
Sample sizes are pre-registered. If owners want continuous monitoring, use anytime-valid confidence sequences (Howard et al., 2021). Repeatedly checking a fixed-n interval is not valid.
6.6 Deterministic behaviour
Both first profiles claim behavior.deterministic: true
(manifest). The claim is
falsifiable:
| ID | Criterion | G1 | G2 | G3 |
|---|---|---|---|---|
| D1 | Byte-identical RFC 8785 canonical findings across repeated runs of the full suite | ≥10 runs (P) | ≥100 runs (P) | ≥100 runs (P) |
| D2 | Identity across GOMAXPROCS ∈ {1, 2, NumCPU}, shuffled check scheduling and concurrent evaluations | 100% (I) | 100% (I) | 100% (I) |
| D3 | Identity across supported architectures (amd64, arm64) | — | report | 100% (I) |
| D4 | Metamorphic: a prefix insertion of k code points shifts every span by exactly k; unrelated segment edits do not change other segments' findings; JSON field relocation keeps findings attached to the field | 100% (I) | 100% (I) | 100% (I) |
| D5 | go test -race and goroutine-leak checks clean | clean (I) | clean (I) | clean (I) |
| D6 | No use of clock, randomness, network or environment in detection (manifest permissions all false) | verified (I) | verified (I) | verified (I) |
Detectors that are not deterministic, such as model-based checks, declare
deterministic: false. They replace D1–D3 with repeated-run variance reporting
and clustered errors (section 12).
6.7 Latency, concurrency and resources
Absolute service SLOs belong to OD-004 (SRE). This document sets only relative and measured-then-frozen rules until those SLOs exist.
| ID | Criterion | G1 | G2 | G3 |
|---|---|---|---|---|
| F1 | Linear scaling: per-KiB cost at manifest max_request_bytes ≤ 2 × per-KiB cost at 64 KiB, including adversarial dense-digit and dense-label inputs (TM-06/TM-07) | ≤2× (P) | ≤2× (P) | ≤2× (P) |
| F2 | Zero timeouts at manifest max_request_bytes and max_concurrency on reference hardware | 0 (I against manifest) | 0 | 0 |
| F3 | p99 detector latency at declared max load (95% order-statistic upper bound, NIST TN 2119 §5.3) | baseline recorded (M) | ≤ 0.5 × configured check timeout_ms (P) | ≤ SRE allocation (R, unset) |
| F4 | Benchmark regression versus last promoted version: benchstat with -count ≥ 10, Mann–Whitney U (benchstat) | report | fail if p < 0.05 and median ns/op or allocs/op worsens > 10% (P) | same (P) |
| F5 | Peak heap/RSS under max concurrency ≤ manifest memory_bytes | report | ≤ (I against manifest) | ≤ (I) |
| F6 | Service-level load test, k6 open model (constant-arrival-rate; avoids coordinated omission), at SRE-approved rate and payload mix | — | report | meets SLO (R, unset) |
Pin hardware class, Go version, GOMAXPROCS and payload corpus digest in every
performance result. Go's regexp package guarantees linear-time matching
(pkg.go.dev/regexp). F1 still applies because label
windows, normalization and post-validation can add super-linear work.
7. Statistical and regression methodology
- Gate on bounds, not point estimates. Every rate-based gate compares a one-sided 95% bound with its target (confidence level: R; 95% proposed). A perfect score on 50 cases only demonstrates a miss rate below about 5.8%.
- Interval method. Use Clopper–Pearson for pass/fail because it guarantees coverage at or above the nominal level. Brown, Cai and DasGupta (2001) recommend Wilson or Jeffreys for estimation, so reports also show Wilson intervals. Do not use Wald intervals.
- Independence.
nis the number of independent base cases (section 5.2). Template families appear as clusters in reports. If a gate set violates the one-entity-per-base-case convention, gate on the base-case indicator or a cluster bootstrap over base cases, whichever is more conservative. - Paired regression testing. A candidate version and the last promoted
version run on identical cases.
- Invariant suites: any newly failing case fails (I).
- Binary per-case outcomes in each critical slice and hard-negative family: exact one-sided McNemar test on discordant pairs (Dietterich, 1998). Fail at p < 0.05 (P). An owner may waive only with a recorded cause, such as an adjudicated label correction.
- Aggregate metrics such as F1: paired bootstrap over base cases (Berg-Kirkpatrick et al., 2012), 10,000 resamples, fixed recorded seed (P).
- Non-inferiority: the 95% upper bound of (baseline recall − candidate recall) must be ≤ δ. Proposed δ = 0.01 at G2 and 0.005 at G3 (R). Apply the same to precision and benign pass.
- Multiplicity. Regression tests on many slices are left unadjusted (P).
This is deliberate: a false alarm costs a review, while a missed regression
costs a leak or an over-block. Absolute-floor gates are per-slice claims. If
owners need simultaneous coverage across k slices, use Bonferroni (α/k) and
recompute
n(R). - Holdout hygiene. Never tune on holdout. Log every holdout evaluation. Any
holdout case inspected for debugging moves to
developmentand is replaced. Rotate part of the holdout each release to limit adaptive overfitting (Dwork et al., 2015). - Prevalence. Corpus precision is not production precision. Report per-opportunity FP rates (per hard-negative case, per benign document). Production precision is estimated only through O1.
- Reproducibility. Every result pins corpus digests, generator seeds, detector and policy digests, evaluator version, Go/toolchain versions, hardware class and raw per-case outputs. Raw outputs are hash-only where values are sensitive. This follows NIST AI RMF MEASURE practice (NIST AI 100-1).
8. Human review and approvals
| Transition | Required review | Approvers |
|---|---|---|
| G0 → G1 | Source citation re-verified at coding time (australian-pii.md); label guideline and conformance matrix reviewed by someone other than the implementer; telemetry fields reviewed for value leakage | Detector owner + independent engineer; privacy reviewer for telemetry |
| G1 → G2 | Holdout and adversarial dispositions reviewed; O1 sample reviewed under an approved privacy process (or signed synthetic-only risk acceptance); OD-010 redaction semantics decided for the category; canary scope and rollback plan | Detector owner, data science, privacy, product; risk approves R values |
| G2 → G3 | Canary results and complaint/override triage; adversarial review (U5); SRE load results (F6); on-error posture (OD-005/OD-006); incident runbook and demotion drill; legal or domain review where the identifier is regulated (section 11) | Risk + product (OD-004), SRE, privacy, legal/domain owner, security |
| Any demotion | Incident review within the SLA that requires approval | Detector owner + risk |
Approvals are recorded against the gate report digest, not a branch or version label. Publish, approve and activate stay separate roles (OD-013).
9. Evidence bundle per promotion
A promotion request contains one immutable bundle:
gate-report.json: per-criterion inputs (n, failures, bound, target, value status, pass/fail), slice table, regression table and waivers with expiry;- gate-definition digest and the profile used;
- detector artefact digest, manifest, source-citation record, policy snapshot digest;
- corpus manifests: digests, partition, licence, generator version and seed, label guideline version, annotator agreement, adjudication log;
- determinism matrix results; adversarial disposition table and declared gaps;
- performance results with environment; benchstat output;
- observed-traffic review summary: counts and reviewer IDs only, no values; or the signed synthetic-only risk acceptance;
- approvals that reference the bundle digest.
The bundle is a separate evidence artefact that points to the package digest, as ADR 0011 requires. Restricted data is never published to make the bundle self-contained.
10. Confidence scores and OD-019
The canonical finding schema requires confidence in [0, 1]. Neither first
detector produces a probabilistic score, and no calibration study exists.
Inventing a threshold such as "block above 0.9" would invent risk policy. It
would also suggest a calibration that does not exist.
Recommendations:
- For G1–G3 of deterministic structured detectors, policies trigger on
category,detector_class,validation_stateand evidence kinds, never on numericconfidence. - Until OD-019 is decided, such detectors emit a documented constant per evidence tuple, for example (shape+checksum) or (shape+exact label). The constant is a label for that tuple, not a probability. Document it in the detector README. It must not change without a minor version bump.
- Contract decision (C-1): add a way to mark scores as uncalibrated, such as
an evidence item
kind: calibrated_scorepresent only when calibrated, or a finding-level flag. Alternatively, makeconfidenceoptional for deterministic detectors. Until then, policy compilation should reject numeric confidence predicates on detectors without a calibration record. - Before any detector, especially a future model-based one, may support a confidence threshold, OD-019 needs this evidence: a calibration set separate from the threshold-selection and test sets; reliability diagrams and ECE with bootstrap CIs per class × category × context slice (Guo et al., 2017); and a threshold chosen on a separate split whose resulting precision and recall meet section 6 bounds on the test split. Calibration must be repeated whenever the detector, the context features or the traffic mix change.
11. First profiles
11.1 HPI-I (au.health.hpi_i, validated_structured)
Source facts: HL7 AU Base au-hpii (16 digits, prefix 800361, Luhn;
invariants inv-hpii-0..2; example 8003619900015717) and the HI Service as the
only authority for issuance
(inventory). The public category
privacy.pii.health_id.au.hpi_i is provisional: the owner adopted it for the
experimental detector (see the
micro-slice record),
but it is not yet approved as public taxonomy. ADR 0010 requires an explicit
mapping.
- Expected outcome: prefix + 16 contiguous digits + Luhn → finding with
validation_state: validated, evidenceshape+checksum. Bare values qualify. The inventory requires only word boundaries and Luhn. - Positive families: bare in prose; labelled (
HPI-I,HPII, full name); FHIR/JSON identifier values; tables, CSV and Markdown; punctuation-adjacent; several per segment, including next to IHI/HPI-O; each protocol path (input, output, tool request, tool response). - Structural hard negatives (H1, invariant): unlabelled
800361+ 10 digits failing Luhn. Labelled Luhn failures (HPI-H5) are excluded from H1 until L-H2 is decided. Also Luhn-valid IHI (800360) and HPI-O (800362), which must be typed as those or not reported, never as HPI-I; 15- and 17-digit runs; 16 digits inside longer digit runs; alphanumeric or underscore adjacency. - Semantic hard negatives (H2): Luhn-valid 16-digit non-health numbers
(payment-card test PANs, order and tracking numbers); digit runs in hashes,
UUIDs, timestamps and decimals;
800361substrings in longer tokens. - Critical slices: protocol path × {bare, labelled, structured field}; output-path (model-generated) slice.
- Label decisions:
- L-H1: are grouped forms (
8003 6199 0001 5717) in scope? - L-H2: with an exact HPI-I label but failing Luhn, emit nothing or
candidate? - L-H3: decimal/float contexts.
- L-H1: are grouped forms (
- Regulatory note: the Healthcare Identifiers Act 2010 regulates use and disclosure of healthcare identifiers (Part 3 Division 4, s 26). HPI-Is are also exchanged legitimately in clinical workflows. Whether a route should block, redact or only record HPI-I is a legal and clinical-product decision. G3 eligibility does not imply it.
11.2 Labelled Victorian Student Number (au.education.vic_vsn, context_bound_structured)
Source facts: VCAA describes the VSN as a randomly generated nine-digit number for
Victorian students under 25. No checksum is published. The Victorian Student
Register is the authority. Part 5.3A of the Education and Training Reform Act
2006 (Vic) governs it, and VCAA states that VSN information may not appear in
shared communications such as class lists
(VCAA).
The public category privacy.pii.education_id.au.vic.vsn (with the au.vic
subdivision level) is provisional: the owner adopted it for the experimental
detector, but it is not yet approved as public taxonomy.
- Expected outcome: exactly 9 digits bound to an approved label or field →
finding with
validation_state: probable, evidenceshape+ (issuer_label|field_path). The label carries the jurisdiction. A bare 9-digit value never produces a finding (C8). - Positive families:
VSN: 123456789,VSN 123456789,VSN#…,Victorian Student Number: …,Victorian Student Number (VSN): …; approved structured-field aliases (vsn,victorianStudentNumber, CSV columnVSN) drawn from the PostgreSQL field-alias table compiled into the snapshot; each protocol path. - Hard negatives:
- bare 9-digit values (C8 suite);
- 9 digits with competing labels (TFN, ACN,
Student ID,student number, other-state student numbers); - the VCAA Student Number (8 digits + letter), which is a different identifier from the same issuer;
- CRN (9 digits + letter); USI;
- VSN label with 8, 10 or alphanumeric values;
- binding failures (label in another sentence, segment, message or field; negation such as "VSN not supplied; ref 123456789"; lists such as "VSN and TFN: …");
VSNused with an unrelated meaning;- JSON keys such as
vsn_hashorvsn_count.
- Critical slices: each label form; structured field versus prose; table/list binding; protocol path; competing-label families.
- Label decisions:
- L-V1: binding scope (same field; same line or sentence; maximum code-point distance; whether the label may follow the value);
- L-V2: is bare
VSNenough, or is education context also required? - L-V3: case sensitivity and punctuation variants;
- L-V4: spaced forms (
123 456 789); VCAA publishes none; - L-V5: table header binding to column values;
- L-V6: cross-segment or cross-message binding. Proposed default: none.
- Domain note: VSN holders are mostly minors. Child-data sensitivity and the statutory sharing rules make legal and education-domain review a G3 approval.
12. Generalizing to other detector classes
| Detector family | Profile basis | Differences |
|---|---|---|
Checksummed validated_structured (IHI, HPI-O, Medicare, ABN) | HPI-I | Type-confusion matrix spans every identifier sharing the structure; checksum near-misses are H1 invariants |
Unchecksummed validated_structured with context gates (CRN, USI, ImmiCard, VIN) | Hybrid | Precision is statistical (H2) with context gates; bare-form behaviour must be an explicit disposition |
context_bound_structured (TFN, driver licence, passport, WWCC, DVA, Ahpra, etc.) | VSN | Jurisdiction becomes a critical slice; C8 bare-value invariant applies; probable ceiling unless an approved authoritative lookup exists (OD-016) |
| Multi-token or reference-backed (address, phone, BSB/account) | Custom | Requires reference-dataset version pinning (OD-017); partial/overlap metrics matter more; strict span still gates redaction |
| Secrets and credentials | Custom | Entropy/prefix families; never live-validate (TM-11); leakage gates identical |
| Model-based safety, jailbreak or prompt-injection checks | Custom | deterministic: false; repeated-run variance and clustered errors replace D1–D3; calibration (section 10) is mandatory before thresholds; attack-catch and benign-pass pairs per the conformance strategy; judge/model pinning; OD-008 for external judges |
Every family keeps the same ladder, legend, statistics, evidence bundle and approval structure. Only metrics, families and dispositions change.
13. CI and reporting integration (proposed)
┌──────────────┐ ┌────────────────────┐ ┌────────────────────┐
│ PR pipeline │──▶│ Nightly / release │──▶│ Promotion request │
│ public data │ │ restricted runner │ │ evidence bundle │
└──────┬───────┘ └─────────┬──────────┘ └─────────┬──────────┘
│ conformance, │ holdout, hard-neg, │ approvals bound
│ determinism quick, │ benign, adversarial, │ to report digest;
│ fuzz smoke, bench │ determinism matrix, │ catalogue lifecycle
│ smoke, gate schema │ benchstat vs promoted, │ + compiler check
▼ ▼ k6 (G3 candidates) ▼
PR status gate-report.json + .md promoted gate level- PR (public data only): conformance and C8 suites, D1 at 10 runs, D4, race, fuzz smoke, benchmark smoke, and validation of the gate-definition schema. Checks run only for detectors whose code, configuration or corpus changed. Failing an invariant blocks merge.
- Nightly/release (restricted runner): holdout, hard negatives, benign and adversarial suites. Logs show only aggregate counts and case IDs. Values and holdout text never appear. Also runs the full determinism matrix, benchstat against the last promoted baseline, and k6 for G3 candidates.
- Gate evaluator (future tool, Python reference harness): reads the gate
definitions plus metrics, computes the bounds itself, and emits
gate-report.jsonand a Markdown summary for the PR or release. It fails on any criterion whose value status isrequires_approvalwithout a recorded approval when evaluating G2/G3. - Wiring: gate JSON under
test/already gets syntax checking fromnpm run validate. Schema validation and the evaluator are follow-up tooling changes, not part of this proposal.
14. Mapping to open decisions
| Decision | What this proposal needs from it |
|---|---|
| OD-004 | Approve the ladder; confidence level (95% proposed); every R floor and ceiling in section 6; non-inferiority δ; risk tier per category × action; SRE latency, availability and coverage-health budgets (F3 at G3, F6, O4) |
| OD-019 | Accept "no numeric confidence thresholds for deterministic structured detectors"; decide contract change C-1; adopt section 10 calibration prerequisites for future detectors |
| OD-005 / OD-006 | On-error posture when a G3 check times out on a blocking route |
| OD-008 / OD-009 | Whether O1 review of traffic is permitted, where, by whom, and evidence retention |
| OD-010 | Redaction semantics per category; defines R1/R3 residual rules |
| OD-013 | Approver role separation for promotion |
| OD-016 | Confirms validated means structural, not HI Service or VSR validated |
15. Decisions needed
Labels: R means risk/product approval; E means engineering may decide.
- R: approve the G0–G3 ladder and its action and lifecycle mapping (architecture + product).
- R (OD-004): approve or replace each proposed G2/G3 floor and the 95% one-sided Clopper–Pearson standard.
- R (OD-019): no confidence thresholds for these detectors; contract change C-1.
- R + legal/domain: whether HPI-I and VSN should ever be blocked, rather than redacted or recorded, on which routes.
- Product + privacy: label decisions L-H1–L-H3, L-V1–L-V6, and
requires_decisiontransformation dispositions (fullwidth digits, invisible characters, homoglyph labels, encodings). - Privacy (OD-009/OD-008): permit O1 traffic review or accept synthetic-only promotion.
- E: grapheme-boundary rule U-1; provisional P values (F1 2×, F3 0.5× timeout, F4 10%, D1 run counts, O2 window, fuzz budgets).
- Architecture + data: public category mapping for the two inventory IDs.
- E: whether to build the gate evaluator and wire schema validation into
npm run validate.
16. Sources
- Hanley & Lippman-Hand, "If nothing goes wrong, is everything all right?", JAMA 249(13), 1983: rule of three.
- Brown, Cai & DasGupta, "Interval estimation for a binomial proportion", Statistical Science 16(2), 2001; NIST TN 2119, Estimating Instrument Performance with Confidence Intervals and Confidence Bounds.
- Miller, "Adding Error Bars to Evals", arXiv:2411.00640, 2024: clustered and paired analysis.
- Dietterich, "Approximate statistical tests for comparing supervised classification learning algorithms", Neural Computation 10(7), 1998.
- Berg-Kirkpatrick, Burkett & Klein, "An empirical investigation of statistical significance in NLP", EMNLP-CoNLL 2012.
- Dwork et al., "The reusable holdout", Science 349(6248), 2015.
- Howard, Ramdas, McAuliffe & Sekhon, "Time-uniform, nonparametric, nonasymptotic confidence sequences", Annals of Statistics 49(2), 2021.
- Segura-Bedmar et al., SemEval-2013 Task 9 evaluation schemas; Stubbs, Kotfila & Uzuner, 2014 i2b2/UTHealth de-identification track, J Biomed Inform 58, 2015.
- Guo et al., "On calibration of modern neural networks", ICML 2017.
- Unicode UAX #15, UAX #29, UTS #39, UTR #36.
- Go
regexplinear-time guarantee;golang.org/x/perf/cmd/benchstat; Grafana k6 open versus closed models. - NIST AI 100-1, AI Risk Management Framework 1.0.
- HL7 Australia AU Base 6.0.0
au-hpii; Australian Digital Health Agency HI Service; Healthcare Identifiers Act 2010 (Cth). - VCAA Victorian Student Number pages; Education and Training Reform Act 2006 (Vic) Part 5.3A.
Revalidate external facts, statutes and tool behaviour before use, as the research index requires.