Detector capability roadmap after the HPI-I/VSN micro-slice
Status: proposal for review. This is planning only and changes no runtime code, contract, or policy.
As of: 3 October 2026
Machine-readable matrix: research/detector-capability-matrix.json
(89 rows). It maps every inventory identifier and every AU.BANK.* capability to a row.
The matrix is a snapshot, not a complete catalogue. New candidates enter as
proposed rows. Rejected rows stay in the matrix so excluded ideas are not
re-proposed without new evidence. IDs, categories, and packages follow
ADR 0010 and the proposed
ADR 0011. They stay provisional
until the ADR 0011 ID grammar and public taxonomy are approved.
Baseline
The approved micro-slice has two detectors: HPI-I (validated_structured) and
exact-labelled VSN (context_bound_structured). They are specified in the
micro-slice record and shipped
as experimental built-ins builtin.pii.au.hpi-i and builtin.pii.au-vic.vsn
(version 0.1.0). Both are disabled by default (G0) and run only under the
observe or redact-experimental profiles of
au-baseline.yaml.
The roadmap assumes the slice delivers three reusable pieces:
- the Luhn validator and boundary rules;
- the label/field binding engine;
- detector-scoped finding IDs and overlap handling.
Ranked next detectors
Each candidate is scored 1–5 on five criteria:
- value: the harm it prevents;
- source: whether an authoritative grammar or checksum exists today;
- precision: how deterministic it is;
- reuse: how much micro-slice machinery it shares;
- readiness: whether any legal, licence, runtime, or governance blocker remains.
Ties go to readiness. Nothing that needs a model runtime or a new contract is ranked.
| # | Detector (runtime ID) | Class / approach | Why now | Gate notes |
|---|---|---|---|---|
| 1 | IHI (builtin.pii.au.ihi) | validated_structured, build | Patient identifier with the highest privacy value; same validator as HPI-I with prefix 800360; already in the taxonomy | Fixture-collision handling (see below); HI Act s 26 legal review before real traffic |
| 2 | Private key blocks (builtin.secrets.private-key) | provider pattern, build | RFC 7468 armour is decisive; most severe secret leak | Needs multi-line spans |
| 3 | Provider credentials batch 1: GitHub, AWS, Stripe, Slack | provider pattern, build | Vendor-published prefixes; GitHub tokens also carry a CRC32 checksum | Never validated live (TM-11); validated for secrets needs an ADR 0010 decision |
| 4 | Invisible/bidi Unicode obfuscation (builtin.security.unicode-obfuscation) | lexical, build | Deterministic defence against tag-character smuggling; also measures the micro-slice's Unicode evasion gaps | Must not fire on emoji ZWJ, Persian/Indic ZWNJ, or legitimate RTL marks |
| 5 | TFN (builtin.pii.au.tfn) | context_bound_structured, build | Very high privacy value; reuses the VSN binding engine | Never validated (ATO algorithm is controlled, OD-016); TFN Rule legal review |
| 6 | Medicare (builtin.pii.au.medicare) | validated_structured with mandatory context gate | Most common Australian health identifier | Pin a licensed primary source first: currently only secondary sources cite the ADHA conformance profile |
| 7 | Payment card PAN (builtin.pii.payment-card) | validated_structured (generic), build | Checksum-backed and PCI-relevant | Exclude 80036[0-2]; redaction format is OD-010 |
| 8 | Email address (builtin.pii.email) | lexical, build | Most requested generic redaction | validation_state semantics for non-identifier PII |
| 9 | BSB + account (builtin.pii.au.bank-account) | context_bound_structured, build | Core banking PII; input to payee-integrity controls later | BSB directory licence before any reference validation |
| 10 | ABN (builtin.pii.au.abn) | validated_structured, build | Authoritative mod-89 check; disambiguates TFN/ACN and PayID-by-ABN | Record only: public business identifier |
Next wave (unranked):
- HPI-O;
- AU phone, IP address, MRZ, and labelled date of birth;
- CRN, USI, ImmiCard, Visa Grant Number, DVA, NSW WWCC, and VIN (label-gated);
- operator canary-token leakage, which generalises the existing
TEST_SECRET_CANARYcheck; - output exfiltration URLs;
- PSPF protective markings, ranked by government-tenant demand;
- JWT, connection strings, and secrets provider batch 2.
IHI evaluation
IHI is ranked first.
Why it ranks first:
- Low engineering cost. It uses the HL7 AU Base profile (CC0): 16 digits,
prefix
800360, and Luhn. The HPI-I validator, boundaries, surface forms, and evidence model carry over unchanged. It belongs in the samehealth-identifiers.yamlfamily. - High harm. It identifies patients and is restricted by HI Act s 26.
- Low false-positive risk. Random 16-digit text passes about 1 in 10⁷ times. Context changes evidence, never state.
The real cost is fixture risk. Tens of millions of IHIs are issued in a space of about 10⁹ Luhn-valid values, so a synthetic checksum-valid IHI has roughly a few per cent chance of being a real patient's number. HPI-I's chance is about 0.1%.
Privacy must approve these mitigations before specification:
- keep the number of distinct values minimal;
- prefer published standard examples;
- never pair a value with names, dates of birth, or providers;
- label every value as synthetic;
- keep values out of telemetry;
- generate the private holdout outside Git.
HPI-O should follow cheaply, but record-only: it is an organisational identifier.
Parallel tracks (not in the ranked list)
Deterministic action/mandate evaluators. These are policy evaluators, not content detectors. They all need a typed, authenticated action envelope, which today's evaluation request and trusted context do not carry.
| Track # | Evaluator | Banking rank |
|---|---|---|
| 1 | Tool/operation allowlist + irreversible-action tier gate | 6 |
| 2 | Tenant/principal isolation | 5 |
| 3 | Tainted-provenance barrier | 2 |
| 4 | Payment intent/parameter binding | 3 |
| 5 | Mandate binding | 1 |
| 6 | Separation of duties | 11 |
The plan places these in Phase 3. The banking research calls them the product wedge, so deciding the envelope ADR earlier is the main lever.
Model-backed detectors (Phase 5). These are integrated, not built:
- prompt-injection/jailbreak classifiers (Prompt Guard 2, ProtectAI v2, Granite Guardian);
- content-safety classifiers;
- person-name NER;
- banking conduct signals: complaint, hardship, scam coercion, advice boundary.
Run an offline benchmark spike in the Python harness now. Production use waits for:
- a runtime decision (OD-007/OD-021);
- model licence approval (the Llama licence is not OSI-permissive);
- self-hosting for sensitive traffic (OD-008).
Conduct and safety classifiers are route_or_friction_only: they can warn, clarify,
or route, but never make the sole adverse decision.
Defer and reject
| Decision | Candidates | Reason |
|---|---|---|
| Defer | Driver licences; WWCC/NDIS/state-student families; government service references; professional register numbers | No authoritative grammar (label-only capture); DVS is OD-016 |
| Defer | Medicare provider number | Only secondary sources; needs an inventory record first |
| Defer | Address | Phase 4 G-NAF/PAF decisions (OD-016/OD-017) |
| Defer | Generic password/entropy secrets; known jailbreak phrases | High false positives; observe-only telemetry at most |
| Defer | Encoded payloads; system-prompt similarity; multi-turn jailbreak | Needs derived-text spans, embeddings (OD-020), or session state (Phase 6) |
| Defer | Hate/harassment enforcement | Dialect false-positive harm; needs governed measurement first (OD-014/OD-015) |
| Defer | Payee integrity; insider purpose; lending evidence; vulnerability | Primarily bank controls or co-design dependent |
| Reject | Inventory exclude identifiers | No safe generic detector. Issuer-private numbers move to tenant-configured labelled-field rules (Phase 3) |
| Reject | Live credential validation | TM-11: testing a credential uses it |
| Reject | Protected-attribute or Indigeneity inference | Prohibited by ADR 0008 |
| Reject | Inline misinformation detection | No reliable detector. Policy-source integrity covers the bounded enterprise case |
Decisions needed
- Approve the ranked order, in particular whether secrets (ranks 2–4) interleave with Australian identifiers.
- Privacy approval for IHI fixture handling.
- Whether ADR 0010's validation state and class semantics extend to non-Australian structured values (payment card, MRZ, checksum tokens), or whether ADR 0010 needs an amendment.
- Medicare source: obtain and licence-check the ADHA conformance material, or keep Medicare context-bound.
- New category roots (
data.,action.,banking.,conduct.) and packages (data-handling,agent-actions,banking-au). - Whether to bring the action-envelope ADR forward.
- The model-runtime path and model licence policy (OD-007/OD-021).
- Package owners for
secrets,security-prompt, anddata-handling(ADR 0011 decision 6).
Maintaining the matrix
- Edit a row's
maturityas it moves fromproposedthroughenforceable. - Add sources before claims: external sources marked
verified_in_this_review: falsemust be revalidated at specification. - Never reuse an ID.
- Once ADR 0011 is approved, specified rows move into
catalog/packages/groundskeeper.dev/*. The matrix then becomes a generated view rather than a source.