ADR-A51: IRI and Identity Policy

Status: Superseded by ADR-A82 Date: 2026-09-23 (amended 2026-09-23 following review) Related: Architecture Review Appendix A, G-05, ADR-A74, ADR-A54, ADR-A68, ADR-A75, docs/architecture/iri-policy.md, docs/architecture/rdf-sparql-patterns-guide.md Drafted by: Agent, autonomous session (P0.1.3). Pending human ratification — see phase-0-status.md. Supersession note: This proposed ADR selected a universal LATTICE identity grammar. It is retained as historical review context only. ADR-A82 replaces that decision with configurable, framework-neutral identity-pattern selection. Amendment note: The initial draft was reviewed in docs/developer/review/ADR-A51-review.md, which found one critical conflict with RDF identity semantics (environment-scoped IRIs), one critical overclaim (uniqueness “by construction” with a truncated hash), and a set of internal contradictions and gaps against the rdf-sparql-patterns-guide.md entity-identity default. The initial disposition claimed those findings resolved; the second review identified remaining gaps. ADR-A82 supersedes this proposed decision rather than extending it further.

Context

data-architecture.md §5.2 currently rejects a graph-family registration when (tenantId, projectId, graphIri) collides on a different hash, family, or owner — a runtime conflict-rejection code path that exists only because identity is not structural. G-05 names the underlying gap: there is no normative IRI minting policy distinguishing a stable lineage identity from an immutable content-addressed revision, and no rule preventing personal data or timestamps leaking into IRIs.

The review additionally established that this policy must be consistent with rdf-sparql-patterns-guide.md, which independently derives an entity-identity default (opaque IRI plus a key-claim registry) and a revision/receipt model (content-addressed graphs plus position-addressed event receipts) from RDF and SPARQL’s own constraints. Where the two disagreed, this revision adopts the guide’s position, because it is derived from the same store-portability and GDPR constraints this ADR is trying to satisfy, and because ratifying two identity models under one platform would be a standing source of bugs.

Decision

Scope

This policy applies to LATTICE-minted IRIs. External vocabulary IRIs are never rewritten by LATTICE. TBox, shapes and vocabulary IRIs (Foundation and every ontology layer) are never tenant- or environment-scoped, and are unaffected by anything in this ADR.

Graph-lineage identity: content and event are two identities, not one

A prior draft used one revision IRI for both “this content state” and “this write happened.” That conflates a content-addressed identity with an event-addressed identity, and produces a cycle if content is ever reverted to a prior state (finding F-9). This ADR separates them:

Kind Form Maps to Mutability
Lineage IRI urn:lattice:{tenantId}:{scope}:{family}:{localName} fnd:PersistentIdentity Stable forever
Content revision IRI {lineageIri}/rev/{schemeVersion}-{contentHash} fnd:Version (a content state) Immutable, content-addressed
Event/promotion IRI {lineageIri}/evt/e{epoch}/{seq} a prov:Activity (a thing that happened) Immutable, append-only, chained by prevRev
Current pointer pat:current triple on the lineage’s version row (normative); an optional materialised {lineageIri}/current graph is a projection, not the identity mechanism — Repointed by a single guarded write at promotion

Registry uniqueness is verified, not merely structural

The original draft claimed identity was “structural by construction” and deleted the registry’s conflict-rejection check entirely. With a hash truncated to 64 bits (semanticHash[0:16] as hex), an adversarial collision costs on the order of 2³² hash operations against attacker-influenced content (Surface contracts, mappings, uploaded ontologies) — not negligible (finding F-2). This ADR corrects both the width and the claim:

Entity identity

Strategy Form Use when Risk if misused
natural-key {base}/{ns}/{urlsafe(keyTuple)} Source has an immutable (not merely “stable”), non-PII business key Key collision across source systems — mitigate with a source-system discriminator; discriminated keys from different sources are a distinct identity until an explicit merge (see below)
derived-hash {base}/{ns}/h/{hex(sha256(schemeVersion\|ns\|enc(keyTuple)))[0:32]} Non-sensitive, immutable composite keys only An unkeyed hash of a low-entropy personal key (email, national ID, phone, MRN) is reversible by dictionary and remains personal data under GDPR (Recital 26) — never use this strategy for sensitive keys
surrogate-claimed {base}/{ns}/s/{uuid4} plus an HMAC-keyed pat:KeyClaim (patterns guide P1) on the natural key Recommended default for entities with mutable or sensitive keys The claim key must be per-tenant and never rotated for existing entities (rotation mints a new scheme version plus an old-to-new alias index, not a re-keyed IRI)
surrogate {base}/{ns}/s/{uuid4} (unclaimed) Genuinely identity-less nodes (a reified span, an extraction candidate) Never idempotent — forbidden for any node re-ingestion must converge onto; declaration and justification required (Rule 6)

{ns} is a registered minting namespace (for example person, order), not an rdf:type assertion. It is chosen once and never renamed — an entity’s classification may change under OWL (multiple types, inferred types, refactored hierarchies), and identity must not depend on it (finding F-10). Readers and tools must never infer rdf:type from IRI structure.

derived-hash was previously described as “avoiding PII in IRIs” for composite or sensitive keys; that was wrong for exactly the sensitive case it named (finding F-6). surrogate-claimed is the strategy for that case, matching rdf-sparql-patterns-guide.md’s recommended default (opaque UUID entity IRI plus a P1 key-claim registry) — this ADR previously stated the opposite as the default and restricted surrogates to identity-less nodes only (finding F-7); both documents now agree.

Key mutability and merges

Key-derived strategies (natural-key, derived-hash) require the key to be immutable, not merely “stable” — business keys are renamed, reissued and merged in practice, and “stable” invited exactly that ambiguity (finding F-8). When two IRIs are later found to denote one entity:

A source-system discriminator (used to avoid cross-source key collision) creates provenance-local identity, not resolved cross-source identity. Cross-source entity resolution is a merge, following the rule above, layered on top of minting.

Rules

  1. IRIs never contain personal data (ADR-A68), including an unkeyed hash of a personal key and any timestamp that could be attributed to a data subject.
  2. IRIs never contain a data version, generation number, or timestamp, except in a content revision IRI’s schemeVersion segment, an event IRI’s epoch/seq segments, or an ADR-A54 infrastructure graph bucket — all of which are identity-level positions, not data-version claims about the resource itself. Lineage and entity IRIs carry none of these.
  3. No IRI carries an environment component. Environments are isolated by dataset and access control (ADR-A54), never by rewriting identity — an IRI denotes the same resource in every dataset it appears in, per RDF’s own semantics (finding F-1). A tenant clone that needs synthetic, distinguishable data uses a distinct tenant (for example a test-fixtures tenant), not a rewritten base.
  4. The tenant segment is an immutable, opaque allocated identifier, never a human-readable name — tenants rename, merge and split, and for single-person tenants a name would itself be personal data (finding F-15). A reserved tenant is allocated for shared reference data used across tenants.
  5. {ns} (entity minting namespace) and {family} (graph lineage family) are registered tokens, never renamed, and never used by readers to infer rdf:type or graph classification (finding F-10).
  6. surrogate (unclaimed) minting in an ingestion plan requires an explicit declaration and justification (G9 — no silent surrogate use). surrogate-claimed does not require this declaration, because it is claim-backed and therefore convergent on re-ingestion.
  7. No @base or relative IRI resolution against a urn: base. RFC 3986 relative resolution against a rootless URN path produces silently wrong results (a reference like <rev/abc> resolving against urn:lattice:...:x does not mean what it looks like it means). Only the fully-written canonical form is ever stored or compared. The canonical lexical form is lowercase (urn:lattice:..., never URN:LATTICE:... — URN lexical equivalence is not RDF term equivalence).
  8. Blank nodes arriving from ingestion are skolemized, deterministically where a natural key exists (natural-key/derived-hash rules apply) and with surrogate-claimed/surrogate otherwise, per the skolem-IRI form in docs/architecture/iri-policy.md §5. Blank nodes are never retained across a request boundary.
  9. Public dereferenceable identity (an https:// mapping for anything published outside the platform, e.g. via LDP/Solid) is not decided by this ADR and must not be assumed. urn: identity is internal-only until a separate decision states otherwise.

Consequences