Architecture decisions
ADR 0003: Unicode code-point offsets
- Status: Accepted
- Date: 2026-10-03
Context
UTF-8 byte offsets are convenient in Go but can split multibyte characters and corrupt redaction. UTF-16 offsets mirror some clients but not Go or JSON text. Grapheme clusters are user-visible units but expensive and unstable for detector interchange.
Decision
Canonical spans use zero-based, half-open Unicode code-point offsets [start,end)
within the exact string identified by content_path. Adapters convert explicitly
to byte or UTF-16 coordinates when required. Text is not silently normalized
between detection and transformation.
Consequences
- Go implementations must map rune indices to byte boundaries before slicing.
- Adapter tests must cover ASCII, emoji, CJK, combining marks, normalization variants, malformed boundaries, overlaps, and adjacent spans.
- Schema descriptions and plugin contracts use the same convention.