Review: RDF operational-patterns documentation set

Review state: Closed 2026-09-23. Dispositioned through the remediation specification IRI-patterns-remediation.md and executed as unit rdf-sparql-patterns-remediation (status). Several fixes were later corrected by iri-patterns-post-3866b21-remediation.

Scope and limits

I reviewed all five documents against each other, against the standards they cite (SPARQL 1.1, SHACL, RFC 9110, RFC 4648, XPath F&O numeric semantics, GDPR Recital 26), and against the rules the documents set for themselves.

Several referenced documents were not provided, so I could not check agreement with them:

I also could not execute the Python examples.

Verdict: the set is not safe to apply to production as written.

The findings below are ordered by severity. Section numbers refer to the patterns guide unless stated otherwise.


1. Blocking defects (production-unsafe)

B1. The epoch guard does not protect against restore (F3 is not actually fixed)

The write path never reads the dataset epoch.

After a restore, every version row still says epoch 3. A stale client holding W/"3-42" therefore matches the restored row. It succeeds against a different history and mints urn:rev:orders/1/e3/0000000000000043. That IRI may already exist in exports, CDC sinks and prevRev references taken before the restore. This is exactly the aliasing F3 and Appendix D.1 claim to have fixed. TCK T-5 would fail.

There is a second hazard. §24.4 says “bump pat:epoch”, but the value being bumped is itself restored data. Suppose you restore twice from the same backup. Epoch 3 is bumped to 4 both times, and epoch 4 is reused. The catalogue (§10.3) already classifies a store-local epoch as unsafe. The guide ignores this.

Fix:

  1. Allocate each new epoch from an external high-water mark (the highest epoch ever issued, plus one).
  2. In the same maintenance transaction, before writers start, rewrite every version row’s pat:epoch to the new value.
  3. Add a dataset-epoch guard to every write (CAS, append, create, delete):
    GRAPH <urn:g:dataset> { <urn:ds:prod> pat:epoch "4"^^xsd:long }   # client's expected epoch == current
    
  4. Adopt one of catalogue §10.3’s fail-safe strategies by name in §24.4.

B2. Lost updates are invisible under the corrected design; the fork query cannot fire

Revision IRIs are deterministic from (aggregate, epoch, seq). So when two writers both win a CAS from 41 under broken isolation, they both mint the same subject, …/e3/0000000000000042, with the same pat:prevRev.

Set semantics then collapse the evidence:

As a result:

This is the §1.1 counter problem again, one level up. The example fork (…0042-b) can never be produced by rev_iri().

Fix:

B3. Revision IRIs collide between the stream form and the aggregate form

Two different version rows mint revisions into the same namespace:

These are two independent counters sharing one IRI namespace, which is F2 reintroduced. The §2.3 and §10.2 examples already show urn:rev:orders/1/e3/… receipts with two different pat:target values. S6 (urn:stream:) and §20.2 (urn:g:) disagree in the same way.

Chapter 22’s claim that “a stream can be appended to … and compare-and-set … without two mechanisms” requires a single version row.

Fix:

B4. ?n + 1 silently changes datatype from xsd:long to xsd:integer

Under SPARQL 1.1 §17.3 and XPath F&O, op:numeric-add on integer-derived types returns xsd:integer. Pinning the input to xsd:long does not prevent this, contrary to the claim in §10.1.

S1 and P3 therefore write counters typed xsd:integer, with three consequences:

Fix: BIND(xsd:long(?n + 1) AS ?n1) in S1 and P3. Add a datatype round-trip test to the TCK.

§14.2 and §19.3 recommend pre-creating pat:seq "0" with no head, so that “the CAS shape above is the only shape ever used”. But §19.1’s WHERE requires pat:head ?prevRev as a mandatory triple. With no head, nothing matches, and every first write returns 412.

Fix: make the head optional:

OPTIONAL { GRAPH <urn:g:meta/17> { <urn:g:orders/1> pat:head ?prevRev } }

Unbound ?prevRev then drops harmlessly from both templates.

B6. The append form does not maintain the head or the chain

S1 increments pat:seq but never updates pat:head and never writes pat:prevRev, yet the §10.2 receipt shows one. On a stream written by both forms:

Fix: in the append update:

B7. Weak ETags never satisfy If-Match

RFC 9110 §13.1.1 requires If-Match to use strong comparison, and a weak tag never matches under strong comparison. With W/"3-41", a compliant server or proxy returns 412 on every conditional PUT. §15.4 and §19.5 (and the Python Version.etag) are therefore non-functional over standard HTTP.

Fix:

B8. Outcome resolution after a timeout is wrong

§19.4 compare_and_set calls resolve() after a TimeoutError and maps ASK == false to CONFLICT. But a false answer during a timeout can simply mean the original request is still executing.

Outcome.UNKNOWN is defined but never returned, which contradicts the Java port (§25.1) and retry rule 3 (§25.5).

Fix: after a timeout, either:

Never map “not yet visible” to CONFLICT.

B9. Retention pruning breaks every quiet aggregate on SHACL-validating stores

§24.2 drops whole month buckets after 400 days, but pat:head on an aggregate untouched for longer than that points at a pruned receipt. On the next write:

Fix, either:

B10. Fencing tokens are checked but never advanced

§7.4 and §16.2 guard with FILTER(?f < mytoken) and never write the new token. This has two effects:

Fix: guard with ?f <= mytoken, and in the same update DELETE ?f / INSERT mytoken.

B11. Global HLC pagination reintroduces the G2 reorder hole

The client stamps the HLC before commit. §21.3’s FILTER(?hlc > last) ORDER BY ?hlc can therefore permanently skip a receipt with a lower HLC that commits after a higher one.

Per-target contiguity (S3) detects the loss only after the consumer has advanced. This is the pre-commit case that §25.7’s stable watermark exists for, but §21.3 does not apply it.

Fix: a global HLC reader must do one of the following:

B12. Contention is sharded for the meta graph only

F12 and T-3 shard the version rows, but every write also inserts into the single graphs urn:g:txn and urn:g:txlog/2026-09. P3 claims all live in urn:g:keys. On an engine with graph- or page-granular conflict detection, every write in the dataset conflicts on these graphs, so A6 is lost anyway.

Fix:

B13. Lock ordering cannot be controlled through the order of WHERE patterns

§10.1 and §19.6 say to “acquire version rows in sorted IRI order” by listing them that way in WHERE. SPARQL gives no control over evaluation or lock order; the optimiser reorders patterns freely.

Fix:


2. Cross-document inconsistencies

C1. The guide still treats superseded ADR-A51 as authoritative

The guide cites A51 as the live authority in:

Fix: repoint each to the relevant section of the catalogue (§6.4, §7.4, §11) or to ADR-A82.

ADR-A51’s header also says “Superseded by ADR-A82” while ADR-A82 is still Proposed. Word it as “to be superseded on ratification of A82”, or record both as proposals.

C2. fnd:replacedBy is used but not defined

P7 prescribes fnd:replacedBy, but the guide’s own list of Foundation terms (§23.3) does not include it. Catalogue §11.4 forbids writing a relation that the selected ontology does not define.

C3. Identifier widths are not fixed, which the catalogue forbids

Catalogue §7.4 and §10.2 both reject variable widths. The guide has