By

Master Data as Applied Epistemology

Two systems disagree about a customer’s address. A survivorship rule quietly picks one, and the workflow moves on. What happened there matters: understanding the quality of that decision requires understanding the philosophy behind it. This post runs dry, but that groundwork is necessary for building real understanding, architecture, and strategy.

Why does that frame matter? A golden record is not a passive mirror of upstream systems; it is a claim. The moment you publish one address, one status, one “true” value, you are not just storing data: you are asserting that this is what the organisation knows and vouching for it. Treating the record as something that knows, rather than something that merely holds a value, is what turns a technical merge decision into an epistemic one. The stakes follow from that shift: if we treat the golden record as neutral plumbing, we stop asking what justifies the value it presents, and we lose the discipline of interrogating our own confidence. Once we accept that the record knows something, we are obliged to ask how it knows it, and that obligation only gets harder to discharge once an AI system, not a rule you can point to, is the one doing the knowing. That is the argument this piece makes.


Every MDM programme eventually must answer a question it never quite phrases this way: how do we know that is true?

Not “where did this value come from,” that’s lineage, and most platforms handle it fine. The harder question is the philosophical one underneath it. When two systems disagree about a customer’s address, and your survivorship rule picks one, you have not just resolved a data conflict. You have made an epistemological claim. You have said: this is what we believe, and here is why we are justified in believing it. Data governance, whether anyone calls it this or not, is applied epistemology with a service-level agreement attached.

Once you see it this way, a lot of the perennial MDM arguments turn out to be old philosophical arguments wearing different clothes.

There is a stronger version of that claim worth naming this is less individual epistemology than social epistemology. A golden record is rarely believed because it is true, it is believed because the organisation has agreed to behave as if it is true. Finance trusts SAP, sales trust the CRM, compliance trusts the KYC system, and when the MDM hub says something different, the real problem is not finding the truth, it is building consensus about what counts as the truth. A golden record is less like a scientific fact and more like a court ruling: it is not accepted because everyone independently verified it, it is accepted because the organisation has agreed on a process for deciding what counts as true. Survivorship logic does not just produce justified beliefs; it acts as an institution that arbitrates disagreement.

Two theories of justification, one architecture diagram

Philosophers have argued for centuries about what justifies a belief. Two answers dominate.

Foundationalism says justification must bottom out somewhere: in basic beliefs that do not need further support, the way a direct sense perception does not need an argument to back it up. Coherentism says there is no bottom, and there does not need to be a belief is justified by fitting consistently into a wider web of other beliefs.

We re-litigate this fight every time we choose a topology. A single authoritative source system for a given attribute (HR owns employee name, the CRM owns customer status) is foundationalism, built into an entity-relationship diagram. Every downstream claim is justified by tracing it back to the designated bedrock. A bank’s KYC platform is a clean case: the government-issued ID number from the identity-verification system is the bedrock field, and every other attribute, name spelling, address, date of birth, only counts as trustworthy to the extent it traces back to that anchor.

Multi-source survivorship is coherentism in production. No one field is foundational. A record is trusted because name, address, birth date, and account number all agree with each other, and disagreement anywhere lowers confidence everywhere. If you have ever sat in a workshop debating whether to trust “the source system of record” or “whichever value has the most corroborating signals,” you were doing three-thousand-year-old philosophy with a Fabric pipeline instead of a chalkboard.

Neither side wins outright, which is why most real architectures are a foundationalist skeleton with coherentist flesh on it.

The pragmatist alternative nobody names

Foundationalism and coherentism do not exhaust the field. MDM practitioners lean on a third theory constantly, without naming it: pragmatism. Nobody calls it that. Everyone does it.

The conversation usually goes something like this: we know source A is right, we know source B is wrong, but the business process only needs 90% accuracy, so we will use source A anyway. That decision is not epistemological, it is economic. It is the pragmatist test for truth, whatever lets the organisation operate successfully, applied without anyone reading a page of philosophy.

Most MDM programmes are not chasing truth. They are chasing operational adequacy. That is not a failure of rigour; it is a rational response to the fact that perfect justification is expensive, and most decisions do not need it. But it does mean a fair amount of what gets labelled data quality is really a cost-benefit judgement wearing an epistemic label.

The Gettier problem, wearing a data quality badge.

In 1963, Edmund Gettier published a three-page paper that broke the tidy definition of knowledge as “justified true belief.” His examples all had the same shape: someone believes something true, and has good reasons for believing it, but the reasons and the truth are connected by pure luck rather than by anything dependable. It does not feel like knowledge. It feels like an accident that happened to land on the right answer.

If you have done a data quality audit, you have met this exact problem without the citation. A customer moved out of an address two years ago, then moved back in last month. The stale record in your legacy system never updated. It still says the correct address today: true, and even “justified” in the narrow sense that the system is a designated source of truth. But the justification and the truth are not connected. The system is right by coincidence, not by tracking reality.

This is precisely why lineage and provenance tracking matter more than point-in-time accuracy checks. A snapshot that is correct today tells you almost nothing if the reasoning behind it is broken. The real question was never just “is this true?” It is “if this is true, is it true for the reason we think it is?” Most data quality frameworks measure the first question obsessively and the second never.

Some master data is discovered; some is invented.

The customer-address examples running through this piece work because addresses have an external ground truth to be right or wrong about. A lot of master data does not have that luxury. Customer status, supplier tier, strategic partner, employee potential, product family: these are not discovered, they are defined. The organisation creates the category by agreeing on it, there is no underlying fact waiting to be found.

That is a real fork in the road. Some master data entities are representational: they describe a world that exists independently of our record-keeping, and our job is to track it accurately. Others are constitutive: they define the thing they describe, and the record-keeping is what makes the category real in the first place. A customer’s date of birth is representational, get it wrong and reality does not change, only our record of it does. Strategic partner is constitutive; a customer becomes one the moment the organisation decides to call them one.

Confusing the two is a common governance failure. We audit constitutive fields for accuracy as if they had a ground truth to check against, when the real question is whether the definition was applied consistently, not whether it was applied correctly.

None of this displaces foundationalism and coherentism, it sits alongside them. Those theories are about discovering a fact that already exists, which is the right lens for a representational field like an address or a date of birth. Constructivism is about creating a fact through agreement, which is the right lens for a constitutive field like strategic partner or supplier tier. An enterprise knowledge system does both at once, often within the same record, and a governance model that only ever asks is these true misses every field where the honest answer is we decided it was.

Explainable to whom? Internalism, externalism, and the audit trail

There’s a related split in epistemology between internalism, which says justification has to be something you can access from the inside (your reasons, laid out and inspectable), and externalism, which says justification can rest on facts about the process itself, like reliability, whether or not anyone can articulate them.

Rules-based matching engines are internalist by construction. Ask why two customer records merged, and the system can hand you the exact rule: matched on national ID and postal code, threshold cleared. A human steward can inspect that reasoning, disagree with it, and override it.

Probabilistic and ML-based matching is externalist by nature. The honest answer to “why did this merge?” is: the model is well-calibrated, and in validation this class of match is correct 97% of the time. That is a real justification. It is just not one a compliance officer can trace like a decision tree, and most governance frameworks were written assuming they would always be able to.

This gap is not a temporary tooling problem. It is the actual philosophical fault line running through every “explainable AI” requirement bolted onto a matching engine after the fact. You are asking an externalist process to produce internalist paperwork.

Then AI shows up and stops asking permission

Everything above describes us justifying our beliefs about data. Bring AI into the pipeline and the question sharpens, though it needs to be asked carefully: an LLM does not have beliefs to justify, it has token probabilities. Calling that coherentism is a useful shorthand, but it is still a shorthand. Philosophically, the model does not justify anything. We do when we choose to act on what it outputs. The accountability layer sits outside the AI, not inside it, which means the real question was never how a non-human reasoner justifies anything, it is why we delegated this decision to a process that cannot. That is a more useful governance question than whether the model can explain itself. The governance challenge is not making AI explain itself, it is explaining why we delegated the decision to AI in the first place, and that’s agent governance, not model interpretability.

Take a concrete case. Feed an LLM-based matching layer two records, “Robert J. Smith, 14 Elm Street” and “Bob Smith, 41 Elm Street”, and ask whether they are the same person. A rules engine fails this match outright: the street number does not line up and there is no exact string overlap on the first name. An LLM, trained on the statistical fact that “Bob” is a common nickname for “Robert” and that transposed digits are a frequent typo, might merge them anyway, and it might even be right. But ask it why, and the honest answer is a probability distribution, not a citation to a rule or a source record.

A large language model has no foundational access to ground truth. It has an extraordinarily dense web of statistical correlation, and any answer it gives is justified only by coherence with that web, not by a traceable link back to an authoritative record. It is coherentism with the floor removed. Ask it to adjudicate a golden record and it will give you something fluent, confident, and structurally ungrounded in the way a compliance audit needs grounding.

That is also the real anatomy of hallucination. It is not the model “lying.” It is a Gettier case manufactured at industrial scale: an output that is sometimes true, sometimes false, generated by a process (statistical likelihood) that is decoupled from the process a steward needs (source authority, provenance, verification). A rules engine is either right for a known reason or wrong in a way you can find. A model can be right for no traceable reason at all, and confident either way.

And AI does not fix bad foundations: it launders them. If the source systems your survivorship rules point to are already stale or wrong, an AI layer sitting on top does not add independent justification. It adds fluency. It takes an unjustified belief and hands it back to you sounding a great deal more authoritative than it has any right to.

The missing ledger: epistemic debt

Ask a steward why a supplier is tiered strategic, or why a matching threshold sits at 0.87 rather than 0.9, and the honest answer is often nobody’s quite sure anymore.

Lineage, provenance, and justification all describe a snapshot: why do we believe this record right now. None of them describes what happens as that justification ages. Call the gap epistemic debt: the accumulated uncertainty created when an organisation keeps relying on a knowledge claim whose original justification has degraded, disappeared, or become unverifiable.

It accrues quietly. The person who owned the source system leaves. The business rule behind a survivorship decision is forgotten. A matching threshold gets tuned once, by someone no longer at the company, and nobody can say why it is set where it is. An AI model gets retrained on new data and the old justification for trusting its output no longer applies to the new one. A supplier classification gets inherited wholesale from a company acquired three reorganisations ago.

In every case, the belief survives and the justification does not. The system keeps working. Nobody remembers why it is right, or whether it still is. Technical debt has a whole vocabulary and a budget line, epistemic debt mostly does not, and it is at least as common a cause of MDM programmes quietly failing.

Which points at the real risk, and it isn’t accuracy?

The dangerous failure mode is not that AI-assisted MDM will be wrong more often than rules-based MDM. It might well be wrong less often. The danger is that it erodes the instinct to ask the question at all.

A raw field value from a legacy system practically invites suspicion: it looks like data, so you interrogate it as data. A synthesized answer from an AI system reads like a conclusion, delivered with the confident fluency of a system that sounds like it already thought it through. A call-centre agent who sees a CRM screen with three conflicting notes about a customer’s account status will hesitate and dig further. The same agent, handed an AI-generated one-line summary that just says “high-value customer, low churn risk”, is far less likely to open the underlying records at all, even though the summary was built from those same conflicting notes. That surface confidence is precisely what Cartesian doubt was designed to cut through, and it is precisely what a good interface makes you forget to apply.

Classical MDM asks us to justify our beliefs about our entities. AI-assisted MDM asks us to justify our trust in a reasoner whose own justifications are frequently inaccessible to us. That is a harder problem and pretending it is the same problem with a faster engine is how governance frameworks quietly stop governing anything.

For decades, MDM has been concerned with the quality of data. AI forces a different question: the quality of justification. Poorly justified knowledge claims have always existed inside enterprises, sitting quietly in stale records and unexamined rules. AI does not remove them, it amplifies them, synthesises them, and presents them with a fluency they never had before. That is only a sharper way of saying what we already know: AI does not fix bad foundations, it launders them. The challenge for governance is no longer managing records. It is managing the evidence by which those records are believed. Stewardship becomes epistemic stewardship, governance becomes justification governance, lineage becomes evidence chains, and trust scores become what they were always meant to be, a quantified account of justification. That is not a rebrand. It is the argument this piece has been making from the first page, followed all the way to its conclusion.

If MDM is applied epistemology, governance must manage justification.

None of this is worth much to a CIO if it stays at the level of metaphor. If master data governance is applied epistemology, the practical question is what governance does differently once it takes that seriously. Concretely, five things change.

Justification lineage. Classical lineage answers where did this value come from. Justification lineage goes one step further and answers why did we trust the source it came from, and does that reason still hold. It is tracked as its own artefact, not inferred from the data pipeline after the fact.

Confidence scores. Every golden record attribute carries an explicit, visible measure of how justified the current value is, not just whether it passed a validation rule. A value can be current and still carry low confidence if the reasoning behind it has thinned out.

Evidence freshness. Confidence decays even when the value does not change. A field last justified three years ago by a source that no longer exists needs a freshness flag, the same way a certificate needs a renewal date, regardless of whether anyone has flagged it as wrong.

Provenance completeness. Not every field needs full traceability, but governance should know, field by field, how much of the reasoning chain is recoverable versus assumed. A gap here is a governance liability whether it has ever caused an error.

Epistemic debt registers. The organisation needs a place to record known justification gaps deliberately, source owner left, threshold untuned, model retrained, so the debt is visible and can be prioritised, instead of sitting invisibly until an audit or an incident surfaces it.

Stewardship responsibilities shift accordingly: a steward’s job stops being just resolve the conflict and starts including record why, and flag it when you cannot. KPIs shift too, from match rate and completeness to something closer to percentage of golden fields with current justification. None of this replaces the existing data quality and lineage tooling, it gives that tooling a job to report against, which is the piece that was missing.

Leave a Reply

About the blog

RAW is a WordPress blog theme design inspired by the Brutalist concepts from the homonymous Architectural movement.

Get updated

Subscribe to our newsletter and receive our very latest news.

← Back

Thank you for your response. ✨

Discover more from The Golden Hour

Subscribe now to keep reading and get access to the full archive.

Continue reading