Conformance and benchmarking strategy
Two distinct test layers
Conformance proves that an implementation obeys the contract: schema handling, span coordinates, coverage states, composition, cancellation, error privacy, and adapter mapping. Benchmarking estimates detection quality, overblocking, robustness, and operational cost. Passing one does not imply passing the other.
Reference harness and case format
Maintain versioned JSONL or Parquet cases and a small independent Python reference evaluator. The same cases drive the Go service and every detector execution mode. Each conformance record includes:
- case ID, schema version, phase, protocol, and synthetic segments;
- expected findings with category, validation state, and code-point spans;
- expected coverage or failure state;
- expected core decision and modifications where policy composition is in scope;
- policy and detector fixture versions;
- tags for locale, adversarial transformation, and risk slice.
Expected values are authored independently, not generated by the Go implementation. The case schema and synthetic examples are the initial interchange format.
Required conformance suites
- JSON Schema/OpenAPI acceptance and rejection.
- Authentication context cannot be overridden by body metadata.
- Stable JSON Pointer extraction and complete Portkey replacement objects.
- Code-point spans across ASCII, emoji, CJK, combining marks, normalization variants, adjacent and overlapping findings, and malformed boundaries.
- Deterministic deny-wins composition and typed on-error handling.
- Required/optional check timeout, failure, unsupported, partial, and cancellation.
- Payload/concurrency/output bounds and no sensitive values in errors or telemetry.
- Plugin manifest/configuration/capability compatibility across execution modes.
- Snapshot digest verification, all-or-nothing activation, and rollback.
Use golden tests for adapters and transformations, fuzz extraction and span
application, property tests for composition invariants, go test -race for
orchestration, and k6 for latency/concurrency/SLO characterization.
Quality metrics
PII and credentials
- exact type-and-span micro/macro precision, recall, and F1;
- overlap/character F1;
- post-redaction leakage and over-redaction;
- results by entity, validation state, locale, and protocol path.
Safety
- precision and recall per policy category;
- action-confusion matrix;
- benign pass rate beside every safety score.
Jailbreak and prompt injection
- attack-catch rate;
- harmful-response attack success rate;
- benign pass rate beside every security result.
Compare versions on identical cases using paired bootstrap 95% confidence intervals and McNemar tests for binary outcomes. Aggregate gains cannot mask a credible regression in a critical category or locale slice.
Corpora
Maintain three partitions:
- public development/practice;
- private fixed holdout;
- rotating incident-derived cases under a separately approved privacy process.
Use synthetic PII and nonfunctional credential canaries by default. Generate deterministic variants for Unicode, zero-width characters, spacing/token splits, encodings, paraphrases/translations, quoted/code/retrieved instructions, role-play, and multi-turn escalation. Pin corpus hashes, licences, seeds, policy digests, detector/model/container versions, judge snapshots, and raw benchmark outputs.
Potential external comparison suites include JailbreakBench, HarmBench, StrongREJECT, AILuminate, XSTest, OR-Bench, garak, PyRIT, and promptfoo. No suite is the product specification. Revalidate every suite and constituent dataset licence; repository licensing may not cover bundled sources or commercial use.
Australian slice
Include synthetic validly structured and hard-negative cases for TFN, Medicare,
IHI, CRN, passport/state licences, +61/04 phones, PO/GPO/Locked Bag and
state/postcode forms, ABN/ACN/public business contacts, invalid checksums, hashes,
invoice/order/ticket references, and Australian person/place-name ambiguity.
First Nations language, names, community content, or identity labels are not routine corpus expansion. Their collection, access, judging, and release require the governance authority in Indigenous Data Governance.
Privacy and corpus register
Treat production examples as personal information until re-identification risk is assessed. Encrypt, minimize, restrict access, set retention, and do not send sensitive or community-controlled material to overseas/cloud judges without explicit privacy, legal, and governance approval.
For every source record version, licence, constituent restrictions, commercial permission, attribution, processing obligations, storage/retention, redistribution, and whether external judges may receive it.
Release gates
Every release needs both:
- absolute quality, benign-pass, latency, and reliability floors by risk tier;
- no statistically credible regression on any critical category or locale slice.
Numerical thresholds and risk tiers are intentionally unresolved in OD-004. Begin in observe mode, measure on approved traffic, tune, and progressively enforce.
The proposed detector ladder (G1 observe, G2 experimental redact, G3 production block) is in Detector maturity gates. It covers corpus minimums, bound-based statistics, approvals and machine-readable gate definitions, and it marks which numbers still need OD-004/OD-019 approval.