Groundskeeper
Architecture decisions

ADR 0003: Unicode code-point offsets

  • Status: Accepted
  • Date: 2026-10-03

Context

UTF-8 byte offsets are convenient in Go but can split multibyte characters and corrupt redaction. UTF-16 offsets mirror some clients but not Go or JSON text. Grapheme clusters are user-visible units but expensive and unstable for detector interchange.

Decision

Canonical spans use zero-based, half-open Unicode code-point offsets [start,end) within the exact string identified by content_path. Adapters convert explicitly to byte or UTF-16 coordinates when required. Text is not silently normalized between detection and transformation.

Consequences

  • Go implementations must map rune indices to byte boundaries before slicing.
  • Adapter tests must cover ASCII, emoji, CJK, combining marks, normalization variants, malformed boundaries, overlaps, and adjacent spans.
  • Schema descriptions and plugin contracts use the same convention.

On this page