Groundskeeper
Research

AI guardrails for consequential Australian banking harm

Research date: 3 October 2026

Scope: Australian banks; customer assistants, employee copilots and tool-using agents across payments, lending, fraud/AML, complaints, hardship, collections, advice and data handling.

Companion inventory: australian-banking-ai-guardrails.json

Executive conclusion

The highest-value gap in common AI safety catalogues is not another classifier for PII, toxicity or prompt injection. It is an action-aware policy layer that can answer, before a consequential step:

  1. Who is acting, for whom, under what authority and purpose?
  2. What evidence influenced the action, and is any of it untrusted, cross-customer, stale or AI-inferred?
  3. Does the proposed action preserve the customer's stated intent, applicable product/process rules and required human accountabilities?
  4. What has happened earlier in the session or case that changes whether this action is safe?

For Groundskeeper, the strongest product wedge is therefore a typed banking action envelope + deterministic policy engine + calibrated semantic detectors + tamper-evident decision evidence. The first controls to build are mandate/authority binding, tainted-action provenance, purpose/tenant boundaries, transaction intent/payee integrity, irreversible-action gates, complaint/hardship recognition and advice/representation boundaries.

These controls complement—but must never replace—bank IAM and mandate stores, payment signing and Confirmation of Payee, fraud/scam engines, AML/CTF and sanctions monitoring, responsible-lending and product-governance systems, core ledgers, records management, operational-resilience controls, or accountable human decision-makers.

Why generic catalogues miss the harm

Generic catalogues inspect text in isolation. Banking harm is usually relational and temporal:

  • a valid instruction is executed for the wrong legal person;
  • individually permitted data fields are combined into an impermissible inference;
  • a clean model response is turned into an unauthorised payment by a tool;
  • a correct policy excerpt is applied to the wrong product, date or customer cohort;
  • a sequence of low-risk steps becomes a high-risk irreversible outcome;
  • a customer's complaint, hardship disclosure or signs of coercion are treated as ordinary service dialogue;
  • a human approval is nominal because the reviewer sees only an AI summary, not source evidence or contradictory facts.

Accordingly, the unit of control should be an event and action sequence, not only a prompt-response pair.

Control classes

ClassMeaningExamplesSafe default
D — deterministic enforcementEvaluates authoritative attributes, explicit policy and typed action parameters. The model cannot waive it.authority scope, tenant boundary, tool allowlist, amount/velocity limit, dual approval, purpose tagdeny, constrain or require approval
P — probabilistic detectorScores ambiguous language, context or behaviour. It must expose confidence and evidence spans.scam coercion, vulnerability, complaint, hardship, advice overreachwarn, clarify, step-up or route; do not silently decide rights
H — human-review triggerRequires a suitably authorised person to decide with source evidence and conflicts visible.financial abuse, disputed authority, adverse lending exception, AML-sensitive explanationpause action and create review packet

The valuable pattern is usually P detects → D constrains → H decides, rather than asking an LLM to enforce policy by prompt.

Action envelope Groundskeeper would need

Every evaluated event should carry, or resolve by opaque reference:

{
  "actor": {"type": "customer|employee|agent", "id": "opaque", "assurance": "..."},
  "principal": {"customer_id": "opaque", "account_id": "opaque"},
  "authority": {"grant_id": "opaque", "scopes": ["..."], "expires_at": "..."},
  "purpose": {"code": "...", "case_id": "opaque", "consent_or_basis": "..."},
  "channel": {"id": "...", "session_id": "...", "trusted": true},
  "provenance": [{"source_id": "...", "trust": "authoritative|untrusted|ai_inferred", "observed_at": "..."}],
  "proposed_action": {"tool": "...", "operation": "...", "parameters": {}, "reversibility": "..."},
  "customer_intent": {"intent_id": "...", "bound_parameters_hash": "..."},
  "policy": {"bundle_id": "...", "effective_at": "..."},
  "sequence": {"step": 0, "prior_decision_ids": ["..."]}
}

Groundskeeper should minimise replicated banking data. Prefer opaque IDs, policy-relevant attributes and signed assertions fetched just in time. Raw identity documents, transaction histories and vulnerability details should remain in the bank's systems of record.

Ranked capability portfolio

Scores are 1 (low) to 5 (high). “Latency” rates suitability for synchronous inline evaluation, not importance. Full scenarios, events, actions, false positives, audit fields and sources are in the companion inventory.

RankCapabilityTypeValueNoveltyFeasibilityLatencyPrimary home
1Mandate, principal and delegated-authority bindingD/H5545Groundskeeper + bank IAM/mandates
2Tainted-action provenance and untrusted-evidence barrierD/P/H5545Groundskeeper
3Transaction intent and parameter-binding integrityD/H5545Groundskeeper + payment signer
4Purpose limitation and use-basis enforcementD/H5435Groundskeeper + privacy/data platform
5Cross-customer/tenant isolation and semantic join controlD/P5445Groundskeeper + retrieval/data layer
6Irreversible-action and blast-radius gateD/H5455Groundskeeper + workflow/tool gateway
7Payee substitution/manipulation controlD/P/H5445Payment/fraud stack, with GK pre-check
8Scam coercion and “safe account” orchestrationP/D/H5434Fraud/scam stack + Groundskeeper
9Complaint, dispute and remediation-rights recognitionP/D/H5444Groundskeeper + IDR case system
10Hardship and collections-state protectionP/D/H5444Groundskeeper + servicing/collections
11Separation of duties and self-approval preventionD/H5355Bank workflow/IAM, enforced at GK gateway
12Advice, representation and product-scope boundaryD/P/H4444Groundskeeper + licensed advice process
13Policy laundering and authoritative-source integrityD/P/H4545Groundskeeper + policy repository
14Record integrity, AI-inference labelling and evidence preservationD/H5444Bank records system + GK evidence
15Sequence-level behavioural policyD/P/H5534Groundskeeper + domain risk engines
16Insider purpose/entitlement misuseD/P/H5334UEBA/IAM + Groundskeeper
17Vulnerability/financial-abuse safe interactionP/D/H5333Groundskeeper + specialist care teams
18Model/tool/supplier change and concentration gateD/P/H4432AI platform, procurement and CPS 230 controls
19AML-sensitive disclosure and tipping-off boundaryD/P/H5434AML case system + Groundskeeper
20Lending evidence and suitability-process integrityD/P/H5433Lending platform, with GK integrity checks

Priority rationale

The top six have three advantages: they prevent harm before execution, rely substantially on deterministic facts, and apply across many domains. Scam and vulnerability detection are vital but inherently noisier and should not become opaque denial mechanisms. Supply-chain controls are strategically important but primarily deployment-time rather than per-token guardrails.

Capability detail

1. Mandate, principal and delegated-authority binding

Scenario. An employee copilot receives “close the account and transfer the balance” from a person recorded as an authority to operate. The authority permits enquiries and bill payments, not closure or transfer to self. A business-banking agent acts on one director's instruction where the mandate requires two.

Control. Resolve the actor, legal principal, account, grant, operation, limits, expiry and required co-signers. Bind them to the proposed tool call. Deny missing or exceeded scope; route disputed capacity, duress or power-of-attorney questions to trained staff.

Why underrepresented. Authentication answers “who logged in”; it does not answer “whose legal power is being exercised for this exact action.” This is a deterministic control, not an LLM confidence score.

2. Tainted-action provenance

Scenario. A copilot reads an emailed invoice or retrieved web page containing instructions to change supplier bank details. The content is legitimate to summarise but must not become authority for a payment or customer-record change. An AI-generated fraud summary incorrectly states that identity verification passed.

Control. Propagate trust labels from each source through retrieval, model output and derived fields. Prohibit high-impact action parameters from being sourced only from untrusted or AI-inferred material. Require an authoritative source or independently re-entered/verified value. Preserve a field-level lineage graph.

Distinct from prompt injection. Prompt-injection detection asks whether content is malicious. Provenance policy remains effective when the content looks benign, the detector misses it, or the issue is stale/wrong evidence rather than an attack.

3. Transaction intent and parameter binding

Scenario. The customer approves “pay ABC Plumbing $1,250 today,” but after tool planning the action is $12,500, a different BSB/account, a recurring payment, or an Osko payment rather than a scheduled transfer.

Control. Canonicalise material parameters at intent capture; display them over a trusted channel; cryptographically bind approval to those parameters; invalidate approval after any material mutation. The model may propose but never silently reinterpret amount, payee, timing, recurrence, account, currency or purpose.

4–5. Purpose, tenant and semantic-join controls

Scenario. A collections copilot uses marketing propensity data; an assistant answers one household member using another member's account notes; an analyst asks an LLM to combine postcode, hardship, merchant and health-adjacent data to infer vulnerability, although no single field is prohibited.

Control. Enforce a declared, permitted purpose and case scope at retrieval and tool-call time. Track subject/customer/tenant labels. Evaluate joins and inferred attributes, not only raw PII. Block or aggregate combinations that exceed approved use; require privacy review for novel inference. These controls operationalise APP 3/6/10 concerns and protect against “mosaic” leakage.

6–8. Irreversibility, payee manipulation and scam coercion

Scenario. An agent creates a new payee and immediately sends a large real-time payment; the customer says a government officer told them to move money to a “safe account,” stay on the phone and not tell bank staff.

Control. Assign every tool operation a reversibility and blast-radius tier. Enforce cooling-off, trusted-channel confirmation, Confirmation of Payee response handling, independent re-authentication and/or human review. A scam/coercion detector can trigger friction, but the payment and fraud platforms own transaction risk and final execution. Do not reveal secret fraud rules or force a victim to confront an abuser.

9–10. Complaint and hardship recognition

Scenario. “You charged me twice and nobody fixed it” is summarised as feedback rather than opened as a complaint. “I am choosing between food and the mortgage” receives budgeting tips while collections continue. An AI summary drops the hardship disclosure, causing a downstream debt-sale or set-off process to proceed.

Control. Use calibrated semantic detectors for explicit and implicit complaint/hardship signals. Deterministically create/preserve a case flag, acknowledgement timestamp, channel and handoff; freeze prohibited automated steps according to bank policy; route to IDR or hardship specialists. Recognition should be recall-oriented, but staff confirm classification and remedy.

11–13. Separation of duties, advice boundary and policy laundering

Scenario. An agent generates a credit exception and approves it using the same service identity. A service assistant uses known balances and goals to recommend a specific investment product while claiming it gave “general information.” A copilot cites an obsolete policy PDF or customer-supplied “policy” to waive a control.

Control. Separate proposer, verifier and actuator identities; reject self-approval and circular multi-agent approval. Classify interaction scope based on substance and data used, not disclaimers. Permit only signed, effective-dated policy sources; surface conflicts and never allow generated summaries to override executable policy.

14–16. Record integrity, sequence controls and insider misuse

Scenario. A model writes an allegation as fact into a customer record, later affecting credit or service. A user stays under every single-action threshold but exports many customer summaries and then emails them externally. An employee invokes a customer copilot without a case or business purpose.

Control. Label observation, customer statement, third-party allegation and AI inference separately; append corrections rather than silently rewriting evidence. Evaluate rolling sequence state—cumulative value, novel payees, privilege changes, retries, overrides, customer count and destinations. Require entitlement and active purpose/case; send behavioural anomalies to existing UEBA/insider-risk systems.

17–20. Vulnerability, supplier risk, AML-sensitive disclosures and lending integrity

Scenario. A customer with limited English, remote access constraints, bereavement or family violence is pressured through a standard workflow; an unannounced model/tool update expands data retention or tool permissions; a copilot tells a customer they are under suspicious-matter review; or a lending assistant invents an expense, suppresses a conflicting document or treats an unverified AI inference as customer-supplied fact.

Control. Detect support needs conservatively and offer accessible alternatives without creating a permanent vulnerability label by default. Gate model/tool versions through inventory, attestation, evaluation and rollback controls. Redact AML investigation state and generate customer-safe explanations from approved reason codes. For lending, bind each material fact to its source, require completion of bank-owned inquiry/verification and suitability gates, surface conflicts, and prevent generative output from becoming an approval or decline reason. AUSTRAC monitoring, SMR decisions, sanctions screening and credit suitability decisions remain in specialist systems.

Enforcement architecture and product boundary

Customer / employee / agent
            │
            ▼
┌──────────────────────────────┐
│ Groundskeeper policy gateway │
│ • schema + provenance        │
│ • deterministic policies     │
│ • semantic detectors         │
│ • sequence state             │
│ • evidence/decision receipt  │
└──────────────┬───────────────┘
               │ allow / constrain / review / deny
               ▼
┌──────────────────────────────┐       ┌──────────────────────────┐
│ Bank workflow / tool gateway │──────▶│ Core banking controls    │
│ typed tools, approvals, SoD  │       │ IAM/mandates, fraud, AML,│
└──────────────────────────────┘       │ sanctions, CoP, lending, │
                                       │ ledger, records, IDR      │
                                       └──────────────────────────┘

Groundskeeper-native

  • action-envelope validation and policy decision API;
  • provenance labels and taint propagation across model/tool steps;
  • purpose, customer/tenant and sequence policy;
  • complaint, hardship, coercion and advice-boundary detectors;
  • signed policy bundle/version checks;
  • human-review packets and tamper-evident decision receipts;
  • replay, shadow evaluation and policy simulation.

Shared integration

  • account-mandate and employee-purpose assertions;
  • Confirmation of Payee and payment-risk responses;
  • customer intent tokens and trusted-channel confirmations;
  • IDR, hardship and specialist-care case creation;
  • model registry, supplier attestations and approved tool manifests.

Upstream/downstream only

Groundskeeper should not become the source of truth for identity, legal authority, customer/account ownership, sanctions lists, suspicious-matter decisions, affordability, target-market determinations, transaction monitoring, ledger state, debt enforcement state, records retention, or final human accountability. It may validate assertions from and route to those controls.

Regulatory and conduct rationale

This report is engineering research, not legal advice. Obligations vary by product, licence, customer and facts. The following sources support the control objectives rather than prescribing a specific Groundskeeper implementation.

  • Operational resilience and security. Current CPS 230 requires effective controls, monitoring, incident/near-miss handling, critical-operation tolerances, scenario testing and management of material service providers and fourth parties; payments, deposit operations and customer enquiries are presumptively critical for ADIs.1 CPS 234 requires information-security capability commensurate with threats and controls over information assets, including those managed by third parties.2
  • AI governance and model risk. APRA's April 2026 letter expects AI inventory, lifecycle accountability, human involvement in high-risk decisions, continuous monitoring, model/tool supply-chain visibility, fallback and exit capability.3 ASIC REP 798 found governance can lag adoption, incomplete inventories and third-party weaknesses, and emphasised fairness, accountability, oversight and consumer risk.4
  • Agentic controls. ASD advises least privilege, tool allowlists, logged tool use, automatic permission restriction, separation of orchestrator/reader/actuator roles, expiring delegation and human approval for high-stakes actions; it locates many enforceable controls in the harness rather than the model.56
  • Privacy and data quality. OAIC guidance treats AI-generated or inferred information about an identifiable person as personal information, highlights reasonable necessity, primary/secondary purpose, transparency, accuracy, human verification and lifecycle assurance.7 Automated-decision transparency amendments commence 10 December 2026 for decisions that significantly affect rights or interests.8
  • Scams and payments. The enacted Scams Prevention Framework requires designated entities to prevent, detect, disrupt, report and respond; most bank obligations commence 31 March 2027, while detailed codes/rules were still being developed at the research date.9 The ePayments Code governs consumer electronic payments, unauthorised transactions and mistaken internet payments.10 These make reliable complaint evidence, intent capture and intervention records important, but do not turn an AI detector into a liability decision-maker.
  • Complaints and hardship. RG 271 contains enforceable IDR standards.11 The 2025 Banking Code incorporates fair service, extra care, financial-difficulty and complaint commitments, including 21-day treatment for certain hardship/default complaints and constraints on debt sale/set-off while hardship is considered.12 ABA guidance calls for identifying hardship signals, specialist authority and tailored care.13
  • Lending, product distribution and advice. Responsible-lending obligations require reasonable inquiries, verification and a not-unsuitable assessment.14 DDO requires products to be designed and distributed for a target market and monitoring of outcomes/significant dealings.15 ASIC's digital-advice guidance applies licensing, advice, competence, monitoring and testing obligations to algorithmic advice.16
  • AML/CTF. AUSTRAC requires ongoing monitoring for unusual transactions/behaviour, documented decisions and appropriate responses, while tipping-off constraints limit what can be disclosed.1718 Groundskeeper can protect the communication boundary but cannot make or replace an SMR, AML/CTF or sanctions decision.
  • Vulnerability and First Nations customers. The Banking Code and ABA vulnerability guidance expect extra care, accessibility, appropriate recording and outcome review.1219 ASIC has highlighted that hardship processes may not fit First Nations kinship/cultural obligations, language and remote-access conditions.20 AIATSIS frames Indigenous Data Sovereignty as the right of Indigenous peoples to govern collection, ownership and application of data about them; this supports governance and engagement, not automated ethnicity-based treatment.21

Australia does not label a single cross-sector “consumer duty” equivalent to the UK's. The practical consumer-outcome expectation is nevertheless visible across efficient/honest/fair licence duties, the Banking Code, DDO outcome monitoring, IDR, vulnerability/hardship guidance, privacy fairness and the SPF. Guardrails should test outcomes and process integrity, not merely disclosure wording.

Phased roadmap

Phase 0 — Contracts and observability (0–8 weeks)

  1. Define the banking action envelope, trust labels, reversibility tiers and decision receipt.
  2. Inventory use cases, tools, data purposes, model versions and downstream control owners.
  3. Make “no authoritative attributes, no consequential action” the default.
  4. Add shadow-only telemetry; establish privacy, retention and access rules for guardrail logs.

Exit: ≥95% of consequential tool calls are typed and attributable to actor, principal, purpose, model/tool version and policy bundle.

Phase 1 — Deterministic prevention (2–4 months)

Build authority binding, tenant isolation, intent/parameter binding, irreversibility tiers, SoD, tool allowlists, policy-source integrity and model/tool version gates. Integrate bank assertions rather than copying source-of-truth data.

Exit: zero bypasses in adversarial tests for wrong-principal, expired mandate, mutated payment, cross-tenant retrieval, self-approval and unapproved tool/version cases.

Phase 2 — Consumer-harm detectors (4–7 months)

Deploy complaint, hardship, coercion, vulnerability and advice-overreach detectors in shadow mode; evaluate by cohort/channel; then use them to clarify, add friction or route—not automatically deny rights. Build safe AML explanation templates.

Exit: agreed minimum recall on severe cases; bounded review workload; no material disparity without documented mitigation; specialist teams accept review packets as usable.

Phase 3 — Sequence and provenance controls (7–10 months)

Add field-level lineage, rolling behavioural state, semantic-join policy, insider-purpose binding, multi-agent delegation chains and signed decision receipts. Integrate fraud/scam, UEBA and records platforms.

Exit: end-to-end scenario tests demonstrate prevention or containment across multi-step attacks, including poisoned evidence and threshold splitting.

Phase 4 — Continuous assurance (10–12+ months)

Automate replay against new models/policies, canary release, drift/outcome monitoring, supplier change attestations, exit/fallback exercises and CPS 230 severe-but-plausible scenarios. Feed incidents, near misses, complaints and overrides back into benchmarks.

Benchmark and evaluation design

Evaluation corpus

Create a versioned, access-controlled suite of synthetic and de-identified cases with independent expected decisions:

  • single-event contrasts: same text, different mandate/purpose/customer/tool risk;
  • minimal pairs: one changed BSB digit, amount decimal, effective date, customer ID or source trust label;
  • multi-turn trajectories: benign steps that become harmful through accumulation, retries, delegation or parameter mutation;
  • counterfactual cohorts/channels: equivalent need expressed in Aboriginal English, plain English, translated dialogue, speech transcript, branch note and terse chat;
  • conflicting evidence: authoritative record versus customer document, stale policy, AI inference and employee assertion;
  • operational failures: model/provider outage, policy service timeout, stale cache, duplicated events, partial tool failure and rollback;
  • human factors: reviewers shown full evidence versus persuasive AI summary, measuring automation bias and decision quality.

Use consumer advocates, financial counsellors, Indigenous and disability representatives, fraud/AML, complaints, hardship, lending, privacy, legal, security and frontline staff in scenario design. Do not infer protected status solely to “balance” a test set; use governed, consented evaluation metadata.

Metrics by control type

ControlPrimary measuresRequired slices
Deterministicpolicy decision correctness, bypass rate, fail-open rate, p95/p99 latency, stale-assertion rejectionchannel, product, tool, authority type, outage mode
Probabilisticsevere-case recall, precision, calibration/Brier score, abstention quality, evidence-span fidelitylanguage/dialect, accessibility mode, channel, scenario severity
Human triggerreview yield, queue time, overturn rate, evidence completeness, reviewer agreement, customer outcometeam, shift, case type, vulnerability/support need
Sequencetrajectory detection lead time, accumulated-harm prevented, state-loss resiliencelong sessions, cross-channel continuation, retries, agent delegation
Governancemodel/policy change escape rate, rollback time, coverage, unresolved near missesvendor/version/use case/critical operation

Report false negatives by harm severity, not only aggregate accuracy. Report false positives by imposed friction: warning, clarification, delay, account restriction or lost access are not equivalent. Latency budgets should reserve synchronous deterministic checks for payment/action paths (typically tens of milliseconds excluding external bank lookups); semantic detectors can run synchronously only where calibrated and operationally justified, otherwise precompute or route asynchronously.

Release gates

  • no high-impact policy may fail open on missing identity, authority, purpose, provenance or version data;
  • no detector alone may make an adverse credit, complaint-entitlement, vulnerability or AML decision;
  • every deny/review outcome must expose stable reason codes, policy/version, material facts and provenance without revealing protected fraud/AML logic;
  • replay must show no regression on all critical deterministic invariants and agreed harm-weighted thresholds;
  • red-team trajectories must cover confused deputy, policy laundering, cross-tenant retrieval, tool substitution, approval mutation, threshold splitting and reviewer manipulation;
  • production rollout is shadow → limited cohort → canary → monitored expansion, with kill switch and tested fallback.

Decisions needed

  1. Product boundary: approve Groundskeeper as policy gateway/evidence layer, not a banking system of record or financial-crime/lending decision engine.
  2. First vertical: choose payments/scams (highest acute loss and inline fit) or complaints/hardship (faster integration and strong conduct value). The recommendation is payments foundation first, while developing complaint/hardship detectors in shadow in parallel.
  3. Data contract: require actor/principal/authority/purpose/provenance/action/sequence fields from integrators; without them, consequential-action controls should fail closed or degrade to no-action mode.
  4. Taxonomy: adopt the proposed AU.BANK.<domain>.<capability> IDs as an extension namespace rather than mixing them into generic “safety” labels.
  5. Assurance ownership: name accountable first- and second-line owners for each policy and require legal/compliance validation before claiming regulatory coverage.

Sources

Footnotes

  1. APRA, CPS 230 Operational Risk Management, current determination in force 1 July 2026. ↩

  2. APRA, CPS 234 Information Security. ↩

  3. APRA, Letter to Industry on Artificial Intelligence, 30 April 2026. ↩

  4. ASIC, REP 798: Beware the gap—Governance arrangements in the face of AI innovation, 29 October 2024. ↩

  5. ASD/ACSC and international partners, Careful adoption of agentic AI services, 2026. ↩

  6. ASD/ACSC, Agentic AI harnesses: the layer above the model, 2026. ↩

  7. OAIC, Guidance on privacy and the use of commercially available AI products, updated 17 January 2025. ↩

  8. OAIC, Consultation on guidance for transparency in automated decision making, 18 May 2026. ↩

  9. ACCC, Scams Prevention Framework, accessed 3 October 2026; see also the Scams Prevention Framework Act 2025. ↩

  10. ASIC, ePayments Code, updated 2 June 2022. ↩

  11. ASIC, RG 271 Internal dispute resolution, 2 September 2021. ↩

  12. Australian Banking Association, 2025 Banking Code of Practice, effective 28 February 2025. ↩ ↩2

  13. Australian Banking Association, Industry guideline: Banks' financial difficulty programs, 1 July 2025. ↩

  14. ASIC, Responsible lending / RG 209. ↩

  15. ASIC, RG 274 Product design and distribution obligations, updated 10 September 2024. ↩

  16. ASIC, RG 255 Providing digital financial product advice to retail clients. ↩

  17. AUSTRAC, Overview of ongoing customer due diligence, accessed 3 October 2026. ↩

  18. AUSTRAC, Tipping off, accessed 3 October 2026. ↩

  19. Australian Banking Association, Customer Vulnerability Guideline, 14 November 2024. ↩

  20. ASIC, Hardship and vulnerable consumers, 2023. ↩

  21. AIATSIS, Delivering Indigenous Data Sovereignty, 2019; see also Indigenous Data Governance in Australia, 2023. ↩

On this page