Persistence approach
Foundational PostgreSQL, stateless evaluation
PostgreSQL is included from the foundation as the authoritative authoring store for correlation knowledge and provenance. The evaluation data plane remains stateless: authenticated deployment configuration selects an in-memory immutable snapshot. Authored policies and schemas live in Git; compiled bundles and detector artefacts move to OCI when distribution is implemented; operational audit goes to an external structured sink.
No request-time policy lookup, dependency resolution, audit write, or general correlation query is required to return a verdict. A compiler reads approved PostgreSQL correlation state, validates it, and emits deterministic snapshot data. Evaluators retain a last-known-good snapshot across database outages.
Initial correlation responsibilities
The foundational schema is responsible for:
- entity/category definitions and validator versions;
- positive and negative context terms;
- known field names, schema aliases, and document-context mappings;
- privacy-safe allowlist and reviewed-finding fingerprints;
- calibration versions and approval status;
- source, licence, purpose, authority, and withdrawal provenance;
- immutable compilation inputs and their snapshot lineage.
It does not store raw prompts, responses, matched PII, secrets, inferred identity, or unapproved community-derived content.
Later control-plane responsibilities
The same PostgreSQL platform may expand to:
- tenants, workspaces, applications, deployments, and authenticated bindings;
- policy publication, approval, activation, rollout, and rollback history;
- detector registrations, capabilities, artefact provenance, and revocation;
- safe audit metadata and search;
- reviewed outcomes and keyed false-positive fingerprints;
- dataset provenance and deletion/withdrawal lineage.
Use transactions, foreign keys, constraints, row-level security, and append-only
history where appropriate. JSONB is for evolving metadata, not a replacement for
relational ownership and lifecycle constraints. The proposed Go stack is pgx
plus sqlc, with Atlas or goose migrations; do not make a large ORM the
architecture boundary.
The compiler/control plane compiles all relevant data into a complete snapshot and pushes it to evaluators. Data-plane installation verifies and atomically swaps the snapshot. Database availability never changes an already installed policy.
Address reference store
G-NAF/address data has a different update, licence, size, and geospatial lifecycle
from tenant control-plane data. Keep it in a distinct database or schema with its
own access role and version marker. Load a complete release, validate counts and
provenance, build PostGIS and pg_trgm indexes, then atomically promote the release.
Do not silently combine partial versions.
Online address candidate retrieval may be justified for an address-specific detector, but it must be explicitly bounded, cached, observable as coverage, and assigned a policy failure posture. Prefer precompiled/cached correlation data for the general hot path.
Correlation data
PostgreSQL can maintain entity definitions, positive and negative context terms,
known field aliases, allowlist fingerprints, calibration versions, and reference
provenance. Full-text search and pg_trgm cover lexical retrieval. Add pgvector
only if benchmarks show structured and lexical methods are inadequate.
Embedding similarity may retrieve candidates; it cannot establish identity, Indigeneity, sensitivity, identifier validity, or a final policy action.
Data minimisation
Do not persist raw prompts, responses, matched PII, or secrets by default. Retain metadata only after its purpose and duration are approved. Production examples are personal information until re-identification risk has been assessed; removing obvious identifiers is not necessarily de-identification.