By

Golden Records Are Beliefs, Not Mirrors

the real shift is from tracking whether our data is right to tracking whether we still know why. This blog post is a structured approach to epistemic governance in master data management, and why it matters more now that AI is doing the matching. It builds on the argument set out in Master Data as Applied Epistemology: that master data governance is applied epistemology, and every golden record is a claim we are justified, or not, in believing. Here we turn that argument into a working framework.

We usually talk about master data governance as a data quality problem is the value correct, complete, and up to date. We want to make the case for a different premise. Every golden record is a governed representation of a proposition the organisation has authorised processes to rely on, and governance is really the discipline of managing what warrants that reliance, and for how long.

In this post, we walk through a structured way to classify our data, choose the right architecture for a given attribute, track five concrete signals most lineage tooling ignores, assess our own maturity, and roll all of this out without trying to do everything at once. We also extend the same thinking to AI-assisted matching and enrichment, where the justification problem gets harder rather than easier.

Executive summary

We argue that master data governance is not primarily a data quality discipline; it is an epistemic discipline. A golden record is not a mirror of reality but a belief the organisation has chosen to trust and act upon. The central governance question therefore becomes: “Why do we believe this value, and is that justification still valid?”

Key ideas

1. Golden records are organisational beliefs

A golden record represents an assertion by the organisation. When finance trusts ERP data or sales trusts CRM data, they are relying on an institutional decision about what should be believed. Survivorship rules and MDM architectures are therefore not neutral technical choices; they embody assumptions about knowledge and trust.

2. Not all master data is the same

  • We distinguish between two dimensions rather than a single binary: referential dependence (how far the value depends on something observable outside the organisation) and institutional dependence (how far the value depends on organisational or legal rules for its meaning). A date of birth scores high on the first and low on the second; strategic partner is the reverse; many real fields score high on both.
  • The governance approach follows the dimension, not a fixed label:
  • Attributes high on referential dependence require accuracy validation.
  • Attributes high on institutional dependence require consistency of definition and application.
  • Attributes high on both require both checks, scoped to the part of the value each one covers.

3. MDM architectures reflect different theories of knowledge

We map common survivorship approaches to philosophical positions:

  • Foundationalism: one authoritative source.
  • Coherentism: multiple corroborating sources.
  • Pragmatism: “good enough” accuracy based on business value.

Our argument is that governance becomes stronger when organisations explicitly acknowledge which justification model, they are using rather than treating architecture as a purely technical decision.

4. Accuracy is not enough

A record can be correct today for the wrong reasons. Using a Gettier-style argument, we claim that governance should focus on whether a belief is properly justified, not simply whether the current value happens to be true. Point-in-time accuracy checks are therefore insufficient.

The five governance mechanisms

Versions of these controls already exist separately across data-quality, MDM, and model-governance practice; the contribution here is organising them around the continuing warrant for reliance on each governed value. The framework groups five governance artefacts:

  • Justification lineage: why was this source trusted, and does that reason still hold?
  • Confidence scores: how justified the value is for a named use, not merely whether it passes validation, and the score is only meaningful once we state what evidence it is built from (source reliability, corroboration, freshness, or completeness).
  • Evidence freshness: has the justification become stale even if the value has not changed?
  • Provenance completeness: how much of the reasoning chain can be reconstructed?
  • Epistemic debt register: a formal record of known justification gaps and assumptions.

The AI governance argument

A significant part of this framework focuses on AI-assisted matching and enrichment.

The core claim is that AI does not solve the justification problem; it makes it harder. It also adds a new claimant to a system built to produce one authoritative answer: an AI-suggested match or enrichment is itself a piece of testimony, competing for space in the golden record alongside human-sourced and system-sourced claims. Large language models operate primarily through statistical coherence rather than direct access to authoritative truth. As a result:

  • AI can produce outputs that appear highly convincing without traceable justification.
  • Hallucinations are framed as industrial-scale “Gettier cases”.
  • A wrong AI-suggested merge that gets accepted stops being a claim and becomes infrastructure: nothing downstream questions it again.
  • The greatest risk is not lower accuracy but reduced scrutiny because fluent summaries discourage users from inspecting evidence.

The memorable phrase is:

“AI does not fix bad foundations. It launders them.”

Accountability for AI

Instead of focusing solely on model explainability, we propose delegation accountability.

For every AI-assisted matching, enrichment, classification, or merge decision, the organisation should maintain a delegation record documenting:

  • what decision was delegated,
  • why it was delegated,
  • evidence that the model is reliable,
  • review cadence and accountable owner,
  • fallback process when confidence drops.

Maturity model

The framework introduces six maturity levels:

LevelDescription
0Unexamined: trust is assumed, not established.
1Traced: lineage exists, but justification is not tracked as a separate artefact.
2Scored: confidence scores exist, but freshness isn’t monitored independently of edits.
3Managed: freshness flags and provenance completeness are tracked; the epistemic debt register is reviewed irregularly.
4Governed: all five mechanisms operate continuously; KPIs are built on justification currency.
5AI-accountable: AI-assisted matching and enrichment carry explicit delegation records, reviewed on the same cadence.

Final takeaway

Our central thesis is that the future of MDM is not about improving data quality metrics alone. It is about managing the quality of justification behind organisational knowledge claims. AI makes this more urgent because it amplifies weakly justified beliefs while making them appear more authoritative. The proposed shift is:

  • From governing data quality to governing justification quality.
  • From stewardship of records to stewardship of knowledge claims.
  • From asking “Is this value correct?” to asking “Do we still know why we believe it?”

What this covers, and what it doesn’t

This approach governs the justification of master data values, and the delegation of matching or enrichment decisions to automated or AI-assisted processes. It does not replace our existing lineage, data quality, or MDM tooling; it just gives that tooling a set of questions to answer that it is not currently being asked.

Every golden record is a claim we’re vouching for

A golden record is not a passive mirror of upstream systems. Publishing one address, one status, one value as the golden version is an assertion: this is what the organisation knows, and the organisation is vouching for it. Once a record is treated as something that knows, rather than something that merely holds a value, governance is obliged to ask how it knows it. To be clear, the record itself has no epistemic agency: it does not know or believe anything. What we mean is that publishing a golden value is the organisation authorising processes to rely on a proposition, and the obligation to justify that reliance falls on us, not on the database.

The obligation has a social dimension as well as an individual one. We rarely believe a golden record because we have independently verified it ourselves; we believe it because we have agreed to behave as if it is true: finance trusts the ERP, sales trust the CRM, compliance trusts the KYC system. Survivorship logic functions as an institution that arbitrates disagreement, closer to a court ruling than a scientific measurement. Selection and justification are not the same thing, and the distinction matters throughout what follows. A survivorship rule selects a value: most recent, longest, preferred source, highest match score. Governance then authorises consumers to rely on that selection. Only sometimes is the selection justified, meaning adequate evidence supports it for the use in question. Treating every golden record as justified by default is the overclaim our epistemic debt argument below is designed to correct.

This is testimony at industrial scale. Once a golden record is published, every downstream system that reads it inherits the claim without re-checking it: thousands of decisions come to rest on one upstream assertion, and the provenance behind that assertion, which source system won the match, which steward approved the merge, what the survivorship rule was, is usually invisible to whoever is consuming the record three systems downstream. The claim survives; the reasons for believing it do not travel with it.

Two consequences follow directly, and the rest of this approach puts them into practice:

  • Every survivorship rule and topology choice is an implicit epistemological position, not a neutral technical default.
  • Justification decays independently of the data value itself, and that decay is easy to miss whenever governance stops at lineage rather than also asking whether the original reason to trust a source still holds.

Not all data is discovered. Some of it, we invent.

  • Before we choose a governance mechanism for a given entity or attribute, we classify it, but a clean discovered-versus-invented split does not survive contact with real fields. Most attributes sit somewhere on two separate dimensions, not in one of two boxes.
  • Referential dependence asks how far the value depends on something observable outside the organisation. A date of birth sits high on this dimension: the underlying event happened whether we recorded it. Strategic partner sits low: there is no external fact to check, only a decision.
  • Institutional dependence asks how far the value depends on organisational or legal rules for its meaning. A “service address” or “legal address” sits high here even though the underlying location exists independently, because which address counts as the service address is a matter of definition, not observation.
  • An attribute can score high on both. A date of birth is mostly referential, but the recognised date can still involve disputed documentation or a legal correction, an institutional question layered on top of a representational one. Strategic partner is mostly institutional, but whether a given supplier satisfies the definition’s criteria is a referential question underneath a constitutive label.
  • Confusing the two dependencies is the common governance failure: auditing the institutional part of a field for accuracy against a fact that does not exist or auditing the referential part only for consistency when there is a fact to get wrong.
  • Here’s the test we use in practice: score each attribute high or low on referential dependence and high or low on institutional dependence separately, rather than forcing it into a single representational-or-constitutive label. A field high on referential dependence gets an accuracy check against an external or authoritative source: this is foundationalism or coherentism at work. A field high on institutional dependence gets a consistency check, was the definition applied the same way every time: this is constructivism at work. A field high on both gets both checks, scoped to the part of the value each one covers.

Putting it into practice

  • For every entity in our data domain map, we score each attribute on referential dependence and institutional dependence separately, rather than tagging it as one or the other.
  • Where an attribute scores high on referential dependence, we apply an accuracy check against an external or authoritative source.
  • Where an attribute scores high on institutional dependence, we apply a consistency check was the definition applied the same way every time, by every steward, regardless of outcome.
  • Where an attribute scores high on both, we apply both checks, scoped to the part of the value each one covers.

Choosing how we justify a value: three architectures

Every choice we make about survivorship architecture re-enacts a centuries-old argument about what justifies a belief, though what follows are heuristic analogies rather than claims that MDM topologies instantiate these philosophical theories in a strict sense: a designated system of record may be authoritative for legal, contractual, or historical reasons that have nothing to do with epistemic foundationalism, and five systems agreeing may simply mean five systems copied the same upstream error rather than corroborating independently. A single authoritative source per attribute is foundationalism, built into an entity-relationship diagram: every downstream claim is justified by tracing it back to one bedrock field. Multi-source survivorship is coherentism in production: no field is foundational, and a record is trusted because several independent fields corroborate each other. A third pattern, used constantly but rarely named, is pragmatism: using a known-imperfect source because the business process only needs a given accuracy threshold, which is an economic decision wearing an epistemic label.

Most of our real architectures are a foundationalist skeleton with coherentist flesh on it, and a pragmatist override on the attributes where perfect justification costs more than the error are worth. The governance value of naming the three explicitly is that it stops us from mislabelling a cost-benefit decision as a data quality decision and losing visibility over it.

Here’s how the three architectures play out in practice:

  • Foundationalism means a single authoritative source per attribute, traced to one bedrock field (a government ID, for instance). It is the right call in a regulated domain with one legally authoritative source, like KYC, identity, or employment records. Overused, it is brittle: if the bedrock source is wrong, there is no cross-check to catch it.
  • Coherentism means multi-source survivorship, where trust comes from mutual corroboration across sources. It works well for attributes with several independent but reconcilable sources, like contact data or account status. Overused, there is no anchor when all the sources drift together, and agreement gets mistaken for accuracy.
  • Pragmatism means using a known-imperfect source because the business process only needs a given accuracy threshold. It suits low-stakes attributes where perfect justification costs more than the error are worth. Overused, cost-benefit decisions get relabelled as data-quality decisions and quietly lose visibility.

Why being right today isn’t the same as knowing

A record can be true today and still not be knowledge, in the same sense a philosopher would recognise from the Gettier problem: a belief that is true, and appears justified, but where the justification and the truth are connected by coincidence rather than by anything dependable. Take a customer who moves out of an address, then moves back in a year later. A stale record that we never updated still shows the correct address today: true, and “justified” in the narrow sense that the system is a designated source of truth, but the justification is not tracking reality; it is a coincidence. This is Gettier-like rather than a strict Gettier case: the value is correct by accident, while the process meant to connect the record to reality has failed.

This is why lineage and provenance tracking need a seat alongside point-in-time accuracy checks, not a blanket priority over them: the two answer different risk questions, and which one should dominate depends on the decision, its consequence, and how reversible an error would be. A snapshot that is correct today tells us almost nothing if the reasoning behind it is broken. The operative question is not is this true, it is: if this is true, is it true for the reason we think it is. Most data quality approaches measure the first question obsessively and the second almost never, which is the gap the five mechanisms below are built to close.

Five things worth tracking beyond accuracy

When we take master data governance seriously as applied epistemology, it changes what we track. None of the five mechanisms below is an invention: versions of confidence scoring, provenance tracking, and exception registers already exist across data-quality, MDM, and model-governance practice. What we are proposing is organising them around a question standard lineage tooling does not ask directly: does the justification for this value still hold?

  • Justification lineage answers why did we trust this source, and does that reason still hold. The artefact is a justification record per attribute, separate from the data pipeline: source, reason for trust, date established, review trigger. The data steward for the domain owns it, and we measure it as the percentage of golden attributes with a current, dated justification record.
  • Confidence scores answer how justified the current value is, not just whether it passed validation, but that only works if we say what the number measures. We treat confidence as use-specific rather than a single property of the value: it is our estimate of how well source reliability, corroboration, freshness, and completeness together support relying on this value for a named use. The same attribute can carry a high confidence score for a marketing use and a low one for a regulatory use, because the evidence bar differs, and the score must be revisited whenever the underlying source or its ownership changes. Without a stated use and a stated basis, a number like 0.91 is decorative precision, not a probability. The artefact is an explicit confidence field on every golden record attribute, tagged to its use, visible in the consuming application rather than buried in a metadata table. The MDM platform owner owns it, and we track the percentage of attributes below the confidence threshold over time.
  • Evidence freshness answers whether the justification has aged even though the value has not changed. The artefact is a freshness flag with a renewal date, triggered independently of whether the value has been edited. The data steward for the domain owns it, and we measure the percentage of attributes past their freshness review date.
  • Provenance completeness answers how much of the reasoning chain is recoverable versus assumed. The artefact is a field-by-field completeness score distinguishing recoverable evidence from assumed or inherited evidence. The data governance lead owns it, and we measure the percentage of fields with fully recoverable provenance.
  • The epistemic debt registers answers where we have deliberately deferred justification, and who knows it. The artefact is a logged register of known justification gaps: a source owner who departed, a threshold that was never tuned, a model that got retrained, a classification inherited from an acquisition. The data governance lead owns it, reviewed by the steering committee, and we track the number of open register entries by age and domain.

What changes for stewards

A steward’s role stops being just resolve the conflict and starts including record why, and flag it when you cannot. Our governance KPIs shift correspondingly from match rate and completeness toward percentage of golden fields with current justification. None of the five mechanisms replaces our existing data quality or lineage tooling; each gives that tooling a job to report against that it does not currently have.

Where AI changes the equation

When we introduce an AI-assisted matching or enrichment step, it sharpens the justification problem rather than solving it. A large language model has no foundational access to ground truth; it has a dense web of statistical correlation, and any output it gives is justified only by coherence with that web, not by a traceable link to an authoritative record. That is coherentism with the floor removed. This is exactly where lineage tooling stops being a compliance nicety and becomes the mechanism by which an AI-augmented MDM system stays falsifiable at all: if a contested match cannot be traced back to its evidence, there is no way to know whether it was ever justified in the first place.

Internalism and externalism in matching

Rules-based matching is internalist by construction: asked why two records merged, the system can hand back the exact rule and threshold, and a steward can inspect and override it. Probabilistic and AI-based matching is externalist by nature: the honest answer to why this merged is that the model is well-calibrated and this class of match is correct at a known validation rate. That is a real justification, but not one that can be traced like a decision tree, and most governance approaches were written assuming they always could. If we bolt an explainable-AI requirement onto a matching engine after the fact, we are, in effect, asking an externalist process to produce internalist paperwork.

Hallucination as an industrial-scale Gettier case

Model hallucination is not best understood as the model lying. It is a Gettier case generated at scale: an output that is sometimes true and sometimes false, produced by a process (statistical likelihood) that is decoupled from the process a steward needs, which is source authority, provenance, and verification. A rules engine is either right for a known reason or wrong in a way that can be found. A model can be right for no traceable reason at all, and it will be equally confident either way.

AI does not fix bad foundations. It launders them.

If the source systems behind a survivorship rule are already stale or wrong, an AI layer on top does not add independent justification. It adds fluency: it takes an unjustified belief and hands it back sounding considerably more authoritative than it has earned. A synthesised summary reads like a conclusion and discourages the scrutiny a raw, visibly conflicting field value would invite precisely the failure mode we need to design our governance against.

The real risk isn’t accuracy, it’s eroded scrutiny

We have been framing AI risk in matching and enrichment as an accuracy problem: will the model get this wrong more often than a rules engine did. That framing understates the risk. An AI-assisted match may well be wrong less often than a rules-based one. What we should worry about instead is that fluency erodes the instinct to check at all.

A steward who opens a CRM record and sees three conflicting notes about a customer’s account status will hesitate and dig further: the friction of visible disagreement invites scrutiny. Hand the same steward a synthesised one-line summary, say high-value customer, low churn risk, built from those same conflicting notes, and they are far less likely to open the underlying records at all. The summary reads like a conclusion rather than a claim, and that is precisely the surface confidence that Cartesian doubt was designed to cut through: a good interface makes us forget to apply it.

This changes what we ask of an AI-assisted matching or enrichment step in production. It is not enough to monitor whether the model’s outputs are accurate; we also need to monitor whether our stewards are still opening the underlying evidence when a synthesised answer looks confident, and design the interface so disagreement stays visible rather than getting smoothed away before a human sees it.

Delegation accountability, not model interpretability

The accountability layer sits outside the AI, not inside it. The governance question that matters are not can the model explain itself, it is why this decision was delegated to a process that cannot justify itself in the way a human steward is required to. Practically, that means we need a delegation record for every AI-assisted matching or enrichment point in production.

A delegation record needs, at minimum:

  • The decision being delegated (match, merge, classification, enrichment)
  • Why it was delegated (volume, speed, a pattern rules cannot capture)
  • The externalist evidence backing reliability (validation accuracy, calibration, review sample size)
  • The review cadence and the accountable human owner
  • The fallback path when confidence falls below threshold

In practice this means running two different trust postures inside the same platform: high-latitude AI suggestion, where the model proposes matches, enrichments, or hierarchy placements liberally, paired with low-latitude human attestation on anything about to become canonical, where a named, accountable person signs off before the claim becomes infrastructure. The two postures must stay visibly distinct. An MDM program that lets AI-proposed data quietly graduate to golden-record status, without a distinguishable attestation step marking the handoff, can keep working operationally while it stops being epistemically sound.

We review delegation records on the same cadence as the epistemic debt register, and treat a model retrain as a trigger event: the old justification for trusting the model’s output does not automatically carry over to the retrained version.

How mature is our governance, really?

We use this model to assess our current governance function and to set a realistic target state. Levels are cumulative: each level assumes the mechanisms from the levels below it is already operating, not replaced.

  • Level 0, unexamined: golden records are treated as neutral plumbing. No justification lineage, no confidence scores. Trust is assumed, not established.
  • Level 1, traced: standard lineage exists (where a value came from), but justification (why it was trusted) is not tracked as a separate artefact.
  • Level 2, scored: confidence scores exist on key attributes, but freshness is not monitored independently of edits, so scores decay silently.
  • Level 3, managed: freshness flags and provenance completeness are tracked. An epistemic debt register exists but is reviewed irregularly.
  • Level 4, governed: all five mechanisms operate continuously. Stewardship KPIs are built on justification currency, not only match rate and completeness.
  • Level 5, AI-accountable: AI-assisted matching and enrichment carry explicit delegation records: what was delegated, why, and what externalist evidence backs the model’s reliability, reviewed on the same cadence as the register.

Rolling this out without trying to do it all at once

We do not implement the five mechanisms and the AI addendum simultaneously. The sequence below classifies first, instruments the highest-risk domains next, and only extends to AI-assisted decisions once the underlying justification mechanisms are already running. Extending delegation accountability to a domain with no justification lineage in place would give us the illusion of AI governance without the substance.

  • Classify: score golden record entities on referential and institutional dependence, and map each to a topology (foundationalist, coherentist, or pragmatist). This gives us a classified data domain map we can use to select the right governance check per field.
  • Instrument: add justification lineage and confidence scoring to the highest-risk domains first, so justification records and confidence fields live in the platform itself, not a side spreadsheet.
  • Monitor: add evidence freshness and provenance completeness tracking on top of the domains we have already instrumented, giving us freshness dashboards and a completeness score per field.
  • Register: stand up the epistemic debt register and assign an owner and review cadence, so it becomes a working register reviewed by the governance steering committee on a fixed schedule.
  • Extend to AI: apply the delegation-accountability addendum to any AI-assisted matching, enrichment, or classification step already in production or planned, producing delegation records for each AI-assisted decision point, reviewed alongside the debt register.

The takeaway

For decades, we have treated master data management as a question of the quality of data. Bringing AI into the pipeline surfaces a different question underneath it: the quality of justification. Poorly justified knowledge claims have always existed inside our enterprises, sitting quietly in stale records and unexamined rules; AI does not introduce them, it amplifies and synthesises them, and presents them with a fluency they never had before.

The practical shift we are asking for is narrow and concrete: stewardship becomes epistemic stewardship, lineage becomes evidence chains, and trust scores become what they were always meant to be: a quantified account of justification, tracked and reviewed with the same discipline we currently reserve for accuracy and completeness.

Leave a Reply

About the blog

RAW is a WordPress blog theme design inspired by the Brutalist concepts from the homonymous Architectural movement.

Get updated

Subscribe to our newsletter and receive our very latest news.

← Back

Thank you for your response. ✨

Discover more from The Golden Hour

Subscribe now to keep reading and get access to the full archive.

Continue reading