Why Persistent Identity Matters

One field, many records, and the question no single system can answer

Every system that touches a farm, a mill, or a watershed keeps its own record, and none of them can say with certainty that the others mean the same thing. This Journal explains why that is a structural problem, and how a permanent identifier separated from every description of the resource addresses it.

RSM ID Learning Center10 min read

Pick any working field in an agricultural valley and count the records that describe it. The county keeps a parcel record under an assessor's number. The farm's management software keeps a field record under its own key, because it needs to plan rotations and log inputs. A crop insurer keeps a location record so it can price risk. A grain buyer knows the field only through the lots that leave it, and a soil scientist knows it as a set of sampling points. Each record is accurate for its purpose, and each was created by someone who had no reason to coordinate with the others. The thesis of this Journal is simple to state and surprisingly hard to act on: until those systems share one permanent way of saying which thing, they cannot reliably cooperate about anything that happens to it.

That is the problem RSM ID was designed for. It is the identity infrastructure of the Regenerative Systems Model, and its central move is to give a resource one permanent identifier — an RID — that is deliberately kept apart from every name, location, owner, and description of the resource. The rest of this Journal explains why that separation, rather than any clever identifier syntax, is what makes identity persist.

The same field, five times

Consider Field 42 in the fictional Willow Creek valley. The county land-records office identifies it as assessor parcel 041-220-17. The farm calls it F-42. The insurer calls it location L-88213. The regional network that buys the grain refers to the farm, not the field, under customer number 44821. Nobody here made a mistake; every identifier does its job inside its own system. The trouble begins the moment a question crosses a boundary: did the cover crop logged against F-42 happen on parcel 041-220-17, and is that the parcel whose soil carbon was measured last spring?

Answering that question today usually means matching. Someone compares names, addresses, boundaries, and dates, and decides that two records probably describe the same field. Matching works often enough to be tempting, but it is fragile in exactly the cases that matter. Boundaries are redrawn, farms change hands, two neighbouring fields share a name, and a spreadsheet column gets renamed. Every match is a judgment, made again by every system that needs it, and the judgments can disagree without anyone noticing.

Figure 1

Rendering diagram…

The diagram is the situation without shared identity: every arrow is dashed because every link is an inference. Nothing in any of the four records says that the others exist.

Why the obvious fixes do not hold

The first instinct is to pick one system and make its identifier the master key. The county number seems authoritative, so everyone could use the assessor's parcel number. This fails for reasons that are worth spelling out, because they recur in every domain. The county's number encodes the county's purpose: when parcels are split or merged for tax reasons, the number changes even though the soil does not. The county also has no interest in identifying the farm, the mill, or the watershed, so the master key covers one kind of resource and leaves the rest unaddressed. And making one participant's key mandatory quietly makes that participant the gatekeeper for everyone else.

The second instinct is to use a web address. A URL is globally unique and easy to share, so why not identify the field by the page that describes it? Because a URL is a location, and locations move. Applications are rebuilt, domains lapse, and content migrates; every one of those ordinary events breaks every reference that pointed at the old address. A URL also conflates the field with one page about the field, and that conflation turns out to be the deeper problem.

The third instinct is to put meaning into the identifier so that people can read it: a code that encodes the county, the farm, and the kind of thing. Readable identifiers feel helpful until the meaning changes. If the identifier says county and the boundary moves, or says farm and the farm is sold, the identifier is now either wrong or must be replaced — and replacing an identifier is precisely the event that persistent identity exists to prevent. The Fieldnote Why an RID should not encode its meaning follows this argument further.

A permanent identifier that says nothing

RSM ID answers all three failures with one design decision. An RID is a registered prefix and an issuer-generated random suffix, such as 451.MYGFVMQE4M6E5HYN, and it identifies one and only one resource referent for the lifetime of the registry. It is never reassigned. It contains no type, place, owner, jurisdiction, or host, so none of those can go stale inside it. It does not belong to the county, the farm, or the insurer, because issuance happens under a registered namespace authority rather than inside any one participant's application.

The word referent matters here. Document 08 defines a resource as any distinguishable entity represented within the RSM universe — digital, physical, conceptual, organizational, ecological, social, or computational (§2.1). A field is a referent. So is the farm that stewards it, the mill that buys its grain, a soil observation taken on it, and the watershed it drains into. Each can receive its own RID, and none of them needs to be publicly visible or retrievable to have one: existence of an identity is kept separate from permission to see it.

With an RID in place, the four records above stop being candidates for matching and become annotations of one identity. The county still keeps its parcel number, the farm still keeps F-42, and the insurer still keeps L-88213. What changes is that each of them can record, explicitly and with its own basis, that its local record refers to Field 42's RID. The resolver demonstration shows exactly this: the parcel's external identifiers are listed beside it, each with the system that asserted it.

Figure 2

Rendering diagram…

The arrows are now solid, and each carries a basis. That detail is not decoration. Document 08 requires that associations with external identifiers preserve the authority and semantics of the external identifier, and that an external identifier is never silently treated as an RSM-issued RID (§7.5). The parcel number remains the county's identifier for the county's purposes; it simply has a durable place to point.

The referent and its representations

The second idea that makes persistence work is the separation between a resource and its representations. A representation is any view of the resource returned when someone asks about it: a landing page for people, a JSON document for software, a JSON-LD description for linked-data tools, a provenance-rich view for an authorized researcher. Representations change constantly. Pages are redesigned, fields are added, access policies tighten. None of that changes which thing is being described.

Document 08 puts it with a deliberately plain example: a physical watershed is not equivalent to the HTML page describing it. The specification's illustrative watershed, 451.6R8T2W9K4M7P0Q5X, may be associated with several legal jurisdictions, a hydrological unit identifier, ecological regions, and stewardship organizations, and every one of those associations can change as boundaries are redrawn or classifications revised (§8.4). The RID does not change, because none of those facts were ever inside it. A researcher and a member of the public may receive different representations of the watershed, but both representations refer to the same watershed, and both remain subject to current disclosure policy.

This separation is what lets the system absorb change without losing track of things. The four records about Field 42 are representations maintained by four parties. Each may evolve on its own schedule. As long as each continues to point at the same RID, a later reader can assemble them into one picture without guessing which descriptions belong together.

Document 08 summarizes the division of labour as four questions that must stay logically distinct (§2.6). An RID answers which resource? An RRN answers under which structured namespace was it named? A resolver answers what authoritative information and representations are currently available? A Semantic Space answers which knowledge and governance environment is responsible for its authoritative record? Persistent identity depends on not letting any one of those answers leak into another.

Following one resource through change

Step 1 of 8

A real-world resource exists

Before any identifier, there is a referent: here, the Wellzai paper Outcome Formation Model.

A resource may be digital, physical, conceptual, organizational, or ecological. It need not be public to have an identity.

Document 08 v1.1 §2.1, §3.6 — an illustrative exampleRSM ID Learning Center · informative material01 / 08

Four frames from the guided Reel: a real-world resource, its registered name, its authoritative space, and an RID that persists while everything around it changes.

The value of a permanent identifier only becomes visible over time, so it helps to follow one resource through a few years of ordinary events. Take the fictional Valley Grain Mill, which buys grain from Willow Creek Farm. When it first joined the regional network, it was registered under the operator's own namespace and named rrn:451:valleygrain:us-ca:valleygrain:facility/cleaning-milling-line. Later, the regional network took over administration of its identity records, and a successor name was registered: rrn:451:willowcreek:us-ca:willowcreek:facility/valley-grain-mill.

In a world without persistent identity, that rename would have broken every contract, invoice reference, and quality record that used the old name. Here it breaks nothing. Both names map to the same RID, the old one is marked superseded rather than deleted, and anyone holding the old name can still resolve it to the mill. You can try the superseded name in the demonstration resolver and watch it map to the same identity.

A year later, the mill decommissions its original cleaning line. That line had its own RID because it was tracked as a distinct facility. Retiring it does not free the identifier for reuse: the identity record stays, a tombstone explains what happened, and the RID will never identify anything else. Meanwhile the mill itself remains active under its permanent RID. The identity lifecycle of each resource is recorded in the registry, separately from the content and publication status of any page describing it.

What persistence asks of people and institutions

None of this works by syntax alone. A random suffix guarantees nothing about persistence; it only removes the reasons an identifier would otherwise need to change. Persistence is an institutional promise, and Document 08 says so directly: the promise must be supported through operational policy and tested recovery procedures, not by the syntactic properties of identifiers (§16.4). Somebody has to run the registry, maintain backups, plan for succession if an operator ceases operation, and refuse to reassign identifiers even when it would be convenient.

Persistence also asks something of every participant: discipline about what an identifier is allowed to mean. Issuing an RID for Field 42 does not make the issuer its owner; the farm, the landowner, and the community stewarding knowledge about the land each keep their own roles, which the Fieldnote Why identity is not ownership examines. And two records that look alike are not thereby the same resource. Document 08 is explicit that similar titles, shared subjects, or common source files do not establish identity equivalence; reconciliation needs evidence and an authorized decision (§15.3). The Fieldnote Why semantic equivalence is not identity equality works through why.

Conclusion: start with the question *which thing?*

The background to Document 08 describes exactly the situation this Journal began with: systems that share concepts and relationships, but whose object identities have historically been bound to individual repositories, paths, and application-specific data models (§1.1). The response is not a bigger shared database. It is a thin, durable layer that lets each system keep its own records while agreeing, permanently, on which thing each record is about.

For a practitioner, the implications are concrete. Reserve a field for the canonical RID beside your existing identifiers rather than replacing them. Treat the RID as opaque and compare it by exact equality. Record external identifiers as associations with a basis, never as substitutes. Keep names, locations, revisions, and descriptions in their own fields, so that each can change without disturbing identity. And reference other organizations' resources by RID, so that your relationships survive their redesigns as well as your own.

The next step is to see how the pieces fit together. The Anatomy of RSM ID explains the RID, RRN, RSN, Identity Record, registries, issuance authority, and resolution as one architecture, and the Viewgraph One referent across many systems shows the Willow Creek example in frames. For the shortest route through the whole idea, follow the Journey of an RSM Identity one frame at a time.