Threat model
Status and method
This is the initial threat model for buffered v1 evaluation. It uses trust boundaries and abuse cases informed by STRIDE, privacy, supply-chain, and ML failure analysis. It is not a completed security assessment. Streaming, a dynamic control plane, external plugins, and production data each require a phase-specific update before release.
Assets and security objectives
| Asset | Objective |
|---|---|
| Prompts, responses, tool arguments/results | confidentiality; bounded, purpose-limited processing |
| Secrets and PII findings | never become a secondary disclosure channel |
| Tenant/deployment identity | authentic, isolated, not body-selectable |
| Policy snapshot and binding | integrity, provenance, atomicity, rollback |
| Detector artefacts/configuration | integrity, least privilege, compatibility |
| Verdict and transformation | integrity, deterministic traceability |
| Audit and evaluation corpora | minimisation, access/retention/lineage controls |
| Service capacity | availability under bounded workloads and dependencies |
| Community-controlled data | authority, purpose, consent, withdrawal, stop-use |
Trust boundaries
Untrusted caller/provider payload
│
▼
┌────────────────────┐ authenticated integration ┌──────────────────────┐
│ Gateway / app PEP │──────────────────────────▶│ HTTP + adapter │
└────────────────────┘ └──────────┬───────────┘
│ canonical event
┌───────────────────────────▼─────────────┐
│ Evaluator + immutable trusted snapshot │
└─────────┬───────────┬───────────────────┘
│ │ future boundaries
audit sink local/remote/WASM detectorPortkey metadata and forwarded headers are adapter-supplied attributes, not trusted tenant identity unless an authenticated allowlisted server-side binding makes them so. A policy block is a successful evaluation; a transport/service failure is not.
Initial threat register
| ID | Threat / abuse case | Required control | Verification |
|---|---|---|---|
| TM-01 | Forge tenant/workspace or request a weaker policy | authenticate integration; resolve server-side binding; body metadata cannot select authority | negative API/contract tests |
| TM-02 | Portkey outage or non-200 silently bypasses enforcement | separate outer failure posture; recommend failOnError:true for enforcing routes; decide OD-005 | outage drills and configuration audit |
| TM-03 | Detector timeout/failure appears clean | explicit coverage states and unjudged; policy-owned required/on-error behavior | timeout/error conformance cases |
| TM-04 | Unicode offsets corrupt or under-redact | code-point offsets; exact source text; checked conversion and deterministic overlap handling | multilingual golden + fuzz tests |
| TM-05 | Raw PII/secrets leak through logs/errors/traces | metadata/hash-only telemetry; structured safe errors; prohibit matched values | sink inspection and canary scans |
| TM-06 | Oversized/decompression/adversarial input exhausts service | body/segment/span limits, bounded fan-out, deadlines, cancellation, no unbounded regex | load/fuzz/adversarial tests |
| TM-07 | Regex/CEL complexity consumes CPU | safe regex engine/limits; typed allowlisted CEL; static/runtime cost and deadlines | compile rejection + worst-case tests |
| TM-08 | Prompt injection manipulates detector or policy | detectors treat content as data; policy not authored by payload; model-based checks isolate instructions and versions | adversarial corpus |
| TM-09 | Malicious overlap causes nondeterministic transformation | stable precedence and merge rules; reject unresolved conflict; return complete gateway replacement | property/golden tests |
| TM-10 | Portkey partial transformedData drops request fields | preserve raw payload; rebuild and return full replacement object | adapter golden fixtures |
| TM-11 | Live credential validation causes use/exfiltration | never test credentials remotely by default; structural/prefix/entropy/context only | egress tests and code review |
| TM-12 | External detector exfiltrates payload | explicit destination/tenant approval, mTLS identity, minimisation, network policy, legal/privacy review | integration policy and egress audit |
| TM-13 | Native plugin compromises host | no Go .so; built-ins trusted; process controls; capability-denied WASM path | packaging/admission tests |
| TM-14 | Plugin resource exhaustion or retry storm | CPU/memory/process/output bounds, deadlines, circuit breakers; retry only declared idempotent calls | fault injection |
| TM-15 | Compromised or substituted artefact | digest pinning, expected signer/provenance, SBOM, trust tier, revocation | tampered artefact tests |
| TM-16 | Partial or malicious policy activation | compile full DAG, typed overrides, approval separation, signature verification, atomic swap, rollback | activation failure tests |
| TM-17 | Tenant weakens mandatory baseline | non-overridable rules and operator-owned settings; compiler rejects forbidden override | compiler tests |
| TM-18 | Policy expression gains I/O or unbounded execution | constrained pure CEL environment, no secrets/I/O/plugins, cost and time bounds | compile/runtime conformance |
| TM-19 | Database outage or tampering changes hot-path behavior | no request-time DB dependency; verified immutable snapshots; least-privilege control plane | outage and stale-snapshot drills |
| TM-20 | Reference-data licence or provenance breach | version/licence register, separate store, attribution/restriction controls, legal gate | dataset admission audit |
| TM-21 | Embedding similarity becomes identity/policy proof | retrieval only; structured evidence and composer retain authority; no identity inference | model/card and policy tests |
| TM-22 | Indigenous data misuse or proxy inference | governance authority, provenance/consent/withdrawal, forbidden proxy purposes, stop-use | governance release gate |
| TM-23 | Aggregate benchmarks conceal local harm | per-category/locale gates, benign pass rate, paired regression statistics | release report checks |
| TM-24 | Replay/duplicate requests produce inconsistent side effects | evaluation is side-effect-free; opaque request IDs; future audits idempotent where needed | replay tests |
| TM-25 | Health/capabilities expose secrets or internal topology | minimal unauthenticated health; sanitize capability details; no configs/paths/values | endpoint security test |
Portkey-specific controls
The initial adapter accepts the documented before/after webhook envelope, maps
beforeRequestHook to input and afterRequestHook to output, and preserves unknown
payload objects needed for full reconstruction. It returns HTTP 200 with
verdict:false for a policy block. Malformed/authentication failures are 4xx;
service failures are 5xx. Portkey's default timeout is approximately three seconds,
so Groundskeeper's total deadline and check budgets must leave gateway margin.
The gateway's timeout/network/non-200 behavior is configured outside Groundskeeper and is not equivalent to an internal detector failure. Enforcement owners must resolve OD-005 and verify actual gateway configuration.
Deferred threat-model work
Before streaming: model incremental UTF-8/SSE decoding, chunk-spanning matches, bounded reassembly, partial disclosure, backpressure, cancellation, tool-call withholding, and transformations after bytes have escaped.
Before external plugins: model supervisor/socket spoofing, local privilege and filesystem/network escape, remote workload identity, tenancy, retries, model/data exfiltration, WASM host calls, cache poisoning, update and revocation races.
Before a control plane: model authorization/RLS, approval separation, confused deputies, trust-root/key rotation, rollout targeting, rollback authorization, audit immutability, deletion, backup/restore, and supply-chain compromise.
Before payload retention or reversible tokenisation: complete a privacy impact and key-management design covering access, purpose, jurisdiction, re-identification, retention, deletion, backups, and breach response.
Residual risk and acceptance
Detectors are probabilistic and bypasses/false positives remain possible. A signed policy proves provenance, not safety. Numerical risk tolerances, SLOs, failure postures, external processing, and retention require named owners in the open-decisions register; engineering defaults do not accept those risks.