Data and documents: master data, document flow and product data
- Джимшер Челидзе
- 17 hours ago
- 18 min read
This article is also available in Russian: Russian version.
These three systems are rarely bought together. Reference data belongs to IT, document flow to legal and administration, product data to engineering. Yet they share one problem: the company does not know where the master copy lives. The same customer is created three times, contracts are approved over email, the bill of materials sits in the lead engineer's head.
Until there is a master copy, everything is built on sand. Accounting counts duplicates, analytics shows three revenue figures, the bot breaks on the first mismatch, AI learns from garbage. That is why the three belong together: they answer one question — where does the company's information live, and who is accountable for it.
We go through each: what it does, what it does not, who needs it, what comes first. At the end — how they connect, the vendors you will meet, a shared checklist.
What's in this article
Master data and reference data: MDM, PIM, the integration layer — one master record per customer and product
Documents: ECM, e-exchange and the contract cycle — approval routes, legally binding exchange, contracts
Product and asset data: PLM, PDM, BIM — bill of materials, versions, the building model
How the three loops connect
The role of AI in data and document work
Vendors you will actually meet
Common mistakes and a selection checklist
Master data and reference data: MDM, PIM, the integration layer
What the class does — and what it does not do
This class exists so that every key object — customer, material, product, piece of equipment — has one master record that every system uses instead of keeping its own copy. Plus the rules: who creates the data, who is accountable for quality, how it flows between systems.
What it does not do: it neither replaces transactional systems nor produces data — it puts order into what others produce. And it is not a one-off. Reference data cleaned once "for the go-live", with no process behind it, is back in chaos within a year.
Subclasses and modules: what the class is made of
MDM — master data management. Master records for customers, suppliers, materials, equipment. The system finds duplicates, merges records and distributes the master version to everything else.
Effect: three versions of one customer become one — cross-system analytics and integrations start working.
Drawback: it runs aground on agreements, not software: departments must agree whose record is the master one, and that is the longest part of the project.
Reference data management. The same foundation one layer down: classifiers, code lists, units of measure. In industry it is a discipline in its own right, often with a dedicated team.
Effect: identical reference data everywhere — data lines up without manual mapping.
Drawback: workarounds bolted on for one system come back at every later integration; reworking costs more than doing it properly first time.
PIM — product information. Content for catalogues and storefronts: attributes, descriptions, images, localisations.
Effect: a product is described once and published to every channel — website, marketplaces, print.
Drawback: it needs a content owner; with nobody accountable for filling it, a PIM is an empty storefront with beautiful architecture.
The integration layer: ESB, iPaaS, API management. The buses and platforms data travels on between systems.
Effect: systems connect through a managed layer, not a cobweb of point-to-point customisations.
Drawback: the layer needs a team; a bus without an owner becomes the bottleneck itself.
Data governance and data catalogues. Who owns the data, who has access, where it lives, what can be trusted. The catalogue is the map.
Effect: data turns from a by-product of systems into a managed asset.
Drawback: it slides into bureaucracy easily — committees and policies with no change in how anyone works with data.
Who needs the class, who is too early — and what has to come first
You need it once the landscape holds two or three systems and cross-system work has started: analytics, integrations, migrations. Symptoms: reports from different systems do not reconcile because of reference data; every integration starts with mapping records by hand; migration is frightening precisely because of the data.
Too early or not needed: with one system a full MDM is overkill — entry discipline and two rules for reference lists are enough. Buying MDM "for growth" before the second and third systems exist means paying to solve a problem you do not have.
Low-code sits on this layer too: building applications over systems that share no reference data and no integration rules just breeds new silos, and in a large landscape the bus plays the role of that foundation.
What it connects to and when to implement. It connects to everything: it takes data from every system and hands the master version back. In time it follows the second or third system. In a trade-and-service trajectory it is the "clean up reference data" step between document flow and ERP; in manufacturing it runs in parallel with the production loop and ERP. Analytics stands on that foundation — as does what we said in the hub about hybrid systems: the heart of a hybrid is a single data model, and this class holds it.
What the class gives the business — and its typical drawbacks
Effect. Cross-system analytics that works — the numbers reconcile for the first time. Cheap integrations: a new system connects to the master record, not a web of copies. Calm migrations: data moves out of one clean source. AI readiness: models learn from consistent data, not three versions of one customer.
Drawbacks people mention less. The effect is invisible while all is fine, which makes the budget the first candidate for cutting: "why are we paying for reference data". The project is more about agreements than technology, and gets stuck where agreements get stuck. And there are no quick wins: the effect accumulates with each new system rather than appearing on go-live day.
How practitioners elsewhere think about this
"Data foundation first" is not a local quirk. DAMA DMBOK, the international body of knowledge on data management, codifies the same disciplines — master data, quality, ownership, catalogues — and is the de facto textbook of the profession: DAMA DMBOK.
Data architecture argues between two approaches. Data fabric — a Gartner term — bets on a technology layer over the sources that connects and enriches data itself. Data mesh — Zhamak Dehghani's concept — decentralises accountability: domains own their data. One of its four principles, "data as a product", requires treating data as a product with consumers and a quality bar, not as residue left by systems — data mesh principles. The point is not the fashionable label but what both share: data acquires owners, an agreed model and measurable quality.
MDM itself comes in styles, light to heavy. Registry: the master holds only references and record matches, systems carry on as before — fast and cheap, but quality in the sources does not change. Consolidation: the master is assembled in a hub for analytics, entry stays in the systems. Centralised: the master is created and edited in the MDM, systems only read — the strongest and most expensive, requiring entry processes to be rebuilt. Start light on one domain and add weight as the organisation is ready, not the reverse.
And on hybrid architecture, covered in the hub. Chris Richardson's database-per-service pattern requires every service to own its data and keep neighbours out of its database: database-per-service. No contradiction with "a single data model": physically data can live anywhere, but reference data, master sources and consistency rules stay shared — otherwise distributed architecture becomes distributed chaos.
What drives the cost. Things you can count before talking to a vendor: how many source systems must be reconciled, how large the catalogue is, how many domains go under governance from day one, whether legacy archives must be migrated and re-verified. Packaged tooling covers most mid-market needs; a full data platform with licences plus services is a different order of magnitude. A pilot takes 3–6 months, a full implementation 9–18. With two or three systems you need no platform at all: a named owner per reference list and a written entry rule cost nothing and remove half the problem. Cost it after a survey.
Documents: ECM, e-exchange and the contract cycle
What the class does — and what it does not do
The document loop covers three things: internal document flow (orders, memos, approvals), legally binding exchange with counterparties and authorities (invoices, delivery notes, acceptance certificates, signed electronically) and the contract cycle. A document stops being paper in a folder and becomes a process with a route, deadlines and an owner.
What it does not do: it manages the document part of processes, not the processes themselves. An ECM will run the approval route for a contract; rebuilding end-to-end procurement is BPM work. The boundary is thin, and companies that outgrow their ECM hit exactly it.
Subclasses and modules: what the class is made of
ECM — internal document flow. Registration, approval routes, task tracking, administrative documents.
Effect: approvals stop living in email and corridor conversations — route, deadlines and the person holding the document are visible.
Drawback: the classic trap is digitising paper bureaucracy as it stands: five paper signatures become five electronic ones and nothing got faster.
Legally binding e-exchange. Invoices, acceptance documents and delivery notes exchanged through e-invoicing networks, filings to authorities, electronic signatures and machine-readable authorisations. The legal frame is national — eIDAS in the EU, the ESIGN Act in the US — and decides which signature type suits which document.
Effect: closing documents arrive in minutes instead of "by the end of the quarter"; couriers and lost originals become history.
Drawback: the effect depends on counterparties — while half are on paper, the company lives in two worlds and pays for both.
CLM — the contract cycle. Drafting from a template, approval, signature, tracking of obligations and dates. Not records management: CLM handles a contract as a deal with obligations, not a document with attributes.
Effect: a contract is prepared from a template in an hour instead of written anew; renewals and obligations are not missed.
Drawback: it needs order in templates and standard clauses — otherwise it replicates the lawyers' inconsistencies at industrial speed.
Electronic archive. Long-term retention with fast search and legal validity over years — where records-retention obligations and, in regulated industries, GxP and FDA 21 CFR Part 11 requirements land.
Effect: a five-year-old document is found in a minute instead of a day in the basement.
Drawback: without retention and disposal rules an archive becomes a digital landfill — everything kept, nothing findable.
Capture and scanning. Recognition of inbound paper: primary documents, correspondence. The bridge to robotisation — see the article on RPA and IDP.
Effect: counterparties' paper gets into the systems without being retyped by hand.
Drawback: recognition needs quality control — more on that in the AI section below.
Who needs the class, who is too early — and what has to come first
Almost everyone needs it: external legally binding exchange is a question of timing, not choice. Internal ECM becomes necessary at roughly 30 people, when approvals start getting lost in email. CLM pays off from a hundred contracts a year; below that, templates and discipline suffice. Symptoms: contracts take weeks to approve; nobody knows where a document is; closing documents are collected in a panic; lawyers work on standard contracts instead of complex ones.
Too early or not needed: for a micro business internal ECM is overkill — email and cloud documents will do. External e-exchange it still needs.
What it connects to and when to implement. Next to it sits the digital workplace — mail, office suite, messenger: the environment document flow lives in. This is an early class: in a trade-and-service trajectory it comes straight after CRM, then reference data, then ERP. It is also an early entry into process management: the approval route is the first process a company can see and measure. Documents tie the landscape together: a contract references a customer in CRM, invoices go to ERP, contract analytics to BI. The rule is the one we started with: straighten the process first, digitise second — otherwise bureaucracy just speeds up.
What the class gives the business — and its typical drawbacks
Effect. Speed: approvals shrink from weeks to days, closing documents from weeks to minutes. Control: you see where a document is and who holds it — execution discipline becomes measurable. Money: less paper, fewer couriers, fewer penalties for missed obligations, fewer lost originals.
Drawbacks people mention less. The main risk is cementing bureaucracy: the system faithfully executes the route you gave it, and redundant approvals live forever once digitised. Then the dual world of the transition: while some counterparties are on paper, costs rise rather than fall — both loops are maintained. And dependence on the top team: if the chief executive signs "in person, in the anteroom", the system dies within a month.
2026: the context
Governments keep widening the perimeter of what must be exchanged electronically, and e-invoicing mandates are the clearest example — the direction of travel is one-way, and "let's wait" only accumulates technical debt. Two constraints shape the choice more than functionality does. Signature law: eIDAS in the EU and the ESIGN Act in the US define what counts as a valid electronic signature, and the answer differs by document type. Residency: GDPR and data residency rules decide where the archive may physically sit, ruling cloud options in or out before anyone opens a feature list. In regulated manufacturing, GxP and FDA 21 CFR Part 11 add audit trail and record integrity requirements a generic ECM does not meet out of the box.
What drives the cost. The number of users and document types you route; whether you configure a cloud ECM or migrate a legacy archive with its retention history; how many integrations the loop must hold. A cloud ECM for a small company is configured in a week; a classic project or migration runs 4–8 months. External e-exchange is costed separately: per outbound document and per signature, inbound usually free. Get a number after a survey.
Product and asset data: PLM, PDM, BIM
What the class does — and what it does not do
PLM (product lifecycle management) answers "what do we make, in which version": bill of materials, documentation, changes, routings — from concept through operation to disposal. PDM (product data management) is its core: storage and versioning of engineering data. BIM (building information modelling) applies the same principle to a construction object: not drawings, but a model of the building with data attached.
What it does not do: it does not design — that happens in CAD, PLM stores and connects the result. It does not plan production: shift schedules and equipment loading are MES. And it does not count money: procurement, inventory and cost are ERP, which takes the bill of materials as given.
Subclasses: what the loop is made of
CAD, CAE, CAM. Where product data is born: geometry, simulation, CNC programs.
Effect: without this layer there is nothing to discuss — everything else works with its output.
Drawback: CAD alone produces files, not order: with no vault the "final version" exists in three copies on three computers.
PDM: engineering data management. One vault of models and documents with versions, access rights, statuses and part-assembly-product links.
Effect: the main source of manufacturing defects disappears — work against an outdated version.
Drawback: it runs aground on discipline, not technology: while part of the team keeps files locally, the vault shows an incomplete picture.
PLM: the full lifecycle. Requirements, design and process engineering, change notices, handover to production, operation.
Effect: a change travels a route rather than correspondence, and every downstream function sees it at once.
Drawback: a heavy project touching design, process engineering, procurement and production at once; run "with IT's own resources" it ends in nothing.
Change and configuration management. Change notices, applicability, a record of which configuration shipped to which customer.
Effect: two years on you can say exactly what a shipped unit is built from — the basis of warranty service.
Drawback: change-notice discipline is inconvenient and reads as bureaucracy, right up to the first expensive warranty claim.
BIM: the information model of the object. A building with its composition, attributes and documents, plus a common data environment shared by designers, client and contractor.
Effect: clashes are found in the model, not on site, and the client receives the object together with its data.
Drawback: the model's value is set by the discipline of populating it; handsome geometry with no attributes is a picture, not an information model.
Who needs the class, who is too early — and what has to come first
You need it when the company develops its own complex products or builds facilities: there is a design function, the product has configurations and versions, changes are regular. Symptoms: the shop floor works from a year-old printout; procurement buys against one bill of materials while assembly uses another; "which version is current" is answered by a person, not a system.
Too early or not needed: if you manufacture to someone else's documentation, or the range is stable for years, an organised file vault with naming rules is enough. A full PLM there buys discipline at too high a price.
What it connects to and when to implement. It comes after discipline in CAD has settled — first everyone designs in one environment under one set of rules, then the vault appears. The product model is primary relative to production: PLM belongs before or alongside MES, because production should receive the bill of materials and routing from a master source rather than reassemble them. Input — CAD; output — routings into MES, the bill of materials into ERP, the equipment record into EAM.
Construction has its own sequence: the common data environment and information management rules appear during design, not after work starts on site.
What the class gives the business — and its typical drawbacks
Effect. One version of the truth about the product, and fewer defects from outdated documentation. Reuse: an engineer finds an existing part instead of drawing a fourth variant of the same bushing, and the item master stops swelling with duplicates. Speed of change: a route, not correspondence. And traceability, without which neither warranty nor claims analysis works.
Drawbacks people mention less. This is the most "human" of the technical classes: it asks qualified specialists to change habits they reasonably believe are good. The project is long and the effect deferred — for six months all anyone sees is extra work tidying the archive. Migrating accumulated data almost always costs more than the licences. And the result depends directly on reference data: a PLM without decent master data produces the same duplicates, only in an organised fashion — which is why it is bound to the master data loop above.
Why this is a data project, not an engineering one
Failure here almost always looks the same: the system is treated as an electronic archive. Files go into it, while the bill of materials, applicability and changes stay where they were — in heads and spreadsheets.
In Dzhimsher Chelidze's "Artificial Intelligence. A Practical Guide to Implementation" this is the third of seven sins of digitalisation: no understanding of what data is and what it is worth. The symptoms are named plainly — data scattered across disconnected spreadsheets, datasets with no owner, quality nobody measures. So is the remedy: appoint owners for the key datasets, run an inventory — what exists, where, in what condition — and introduce quality metrics.
For PLM that becomes three questions to answer before choosing a system. Who owns the bill of materials — not "the department", a name. What share of the item master is duplicated. And by what rule a version becomes effective. Until there are answers, any system will carefully store the mess.
There is a harder formulation of the running order too. In the same book the project lifecycle puts data diagnostics second, and gives it the right to stop the project. A red flag on data quality means "stop", not "let's start and sort it out as we go". Here that is the most useful stage of all, because the state of the archive is usually worse than expected — better learned before the contract than after.
2026: the context
Two processes run at once. In engineering companies, consolidation onto fewer CAD and PLM platforms continues, and the obstacle is rarely functionality. It is migrating data accumulated over decades and retraining designers who have worked one way for twenty years. In construction the pressure comes from information management requirements: ISO 19650 gives the common language for how information is produced, exchanged and handed over across the asset lifecycle, and a growing number of public and institutional clients write those deliverables into the contract. There BIM is no longer a maturity question but an admission requirement — you either produce the model and its data, or you do not qualify to bid.
What drives the cost. A pilot takes 2–3 months; a "quick start" for one design group is shorter, 2–6 weeks, with narrower coverage. Licences per seat are rarely the main line: the budget goes on surveying the design office, cleaning up and migrating the legacy archive, the classification work behind the item master, and the process redesign around change notices. Full-scale projects are costed after that survey; any number before it is a guess.
How the three loops connect
The order between them is not arbitrary, and matters more than the choice of any specific system.
Reference data first. Not as a platform, but as an agreement: who is the master source for a customer, a material, a piece of equipment. That needs no software at all, and removes half the future problems in the other two loops. Document flow on dirty reference data replicates inconsistent counterparty details; PLM on dirty reference data reproduces item duplicates in an organised way.
Document flow — early, and for almost everyone. It needs no mature landscape and shows a fast, visible effect: a contract stops sitting in somebody's inbox for three weeks. For a trade-and-service company it is one of the first steps after the customer loop.
Product data — wherever there is something to design. Machinery, instruments, construction. The master copy is born before the production systems: without a bill of materials and a routing, the production loop is fed by hand.
Where all this is handed on. Reference data — to every transactional system and to analytics. Contracts and primary documents — to the accounting and planning loop. Bills of materials and routings — to the production loop. And the whole layer is the foundation for routine automation and AI: a bot works exactly as well as the data underneath it is organised.
The rule common to all three: agreement first, system second. None of them creates order — they lock it in and defend it from erosion.
The role of AI in data and document work
The data layer is one of the few places where AI pays back almost at once: the tasks are high-volume, repetitive and checkable.
What already works. Record matching and deduplication: a model sees that "Acme Ltd." and "Acme Limited" are one customer, more accurately than hand-written rules. Extraction from scans and files: numbers, amounts, parties, dates filled in without a clerk. Classification and routing of inbound documents. Contract review against a checklist with deviations from standard clauses highlighted. Natural-language search across the archive. In the engineering loop — geometric search for similar parts, a direct hit on item duplicates, and recognition of paper drawings during archive conversion.
What is coming (a forecast, not a fact): data steward agents that find a problem, prepare the correction and hand it to a human to confirm; contract assistants drafting the schedule of objections; a first-cut routing generated from the 3D model.
Conditions without which none of it flies. Data: an agreed model and common templates — AI speeds up the clean-up but does not replace the agreement on whose record is the master one. Process: extraction accuracy is not a property but a process; without sampling checks it degrades quietly. People: every reference list and route has an owner, or nobody acts on what the model finds.
The honest boundary and the standard failure. A company buys AI-driven cleansing of its reference data instead of fixing the entry process — six months later the data is dirty again. In the same book this is trap No. 3 of six, "AI as a sticking plaster": solving an organisational problem with a technical one. The rule: fix the process first, automate second.
Vendors you will actually meet
These are names you will meet on shortlists, not a ranking. Vendors carry several products and editions with different scope, so check which one covers your case before comparing.
MDM, reference data, governance and catalogues — Informatica · Collibra · Alation · Ataccama · SAP Master Data Governance · Stibo Systems
PIM and product content — Akeneo · Salsify · inRiver · Stibo Systems
Integration layer (ESB, iPaaS, API management) — no short list: the segment is fragmented, and the bus is more often assembled from platform components than bought as a product
ECM and document management — OpenText · M-Files · SharePoint · Box · DocuWare · Laserfiche
Contracts and e-signature — DocuSign · Adobe Acrobat Sign · Icertis · Ironclad
CAD, PLM and PDM for engineering — Siemens Teamcenter · PTC Windchill · Dassault ENOVIA · Autodesk Vault
BIM and construction information management — Autodesk Construction Cloud · Bentley · Trimble
Two rows deserve a caveat. The integration layer is where "buy a product" misleads most: the bus usually ends up assembled from components you already own, and the decision is who runs it, not whose logo is on it. And PIM tasks are covered well enough by modules inside commerce platforms that a separate PIM earns its place only once catalogue and channel count outgrow them.
The e-exchange network is dictated by counterparties and local rules: connect where most of yours already are, and check what signature type your documents require. In the engineering loop the choice starts with industry specifics and the volume of legacy data — switching costs are set not by licences but by what must be migrated and re-verified.
Common mistakes and a selection checklist
The mistakes repeat across all three loops, so here they are together.
Digitising the mess as it stands. Five paper signatures became five electronic ones; the reference list moved to the new system with its duplicates. Not faster, just more expensive.
Implementing before the agreements exist. Until departments agree whose record is the master one and who owns the route, the system only records the conflict.
Starting without an inventory. Archives and reference data are always worse than expected. Finding out after signing is expensive.
Running it as an IT project. Without lawyers, designers, procurement and records staff, the software changes, not the way work is done.
Letting the dual world drag on. With no plan for moving counterparties and departments across, paper and electronic loops run side by side for years.
Bolting workarounds onto reference data. A customisation for one system comes back at every later integration.
Dropping control over recognition. AI extraction accuracy degrades unnoticed if it is not measured on samples.
Hoping AI will clean it up by itself. It will — and again six months later, until the entry process is fixed.
Checklist before choosing a system
Owners are named: every key reference list, every document type, the bill of materials — by name, not "the department".
An inventory is done: how many records, what share are duplicates, what condition the archive is in. Numbers on paper, not "roughly fine".
Routes and entry rules are written down — before automation, not after.
Integration with the transactional loop is checked: how the master record reaches the systems that use it, and what happens on a mismatch.
Ongoing support is budgeted: who cleans, who approves changes, how often. I would walk away from a project with that line empty — it returns to its original state within a year.
What next
The overall map of classes is in Enterprise IT systems: the map. Other guides in the series: Production, assets and warehouse · Customers and sales · Money, planning and analytics · Routine automation and AI · People and the IT function · IT infrastructure.
To go deeper, see Dzhimsher Chelidze's books "Digital Transformation for Directors and Owners" and "Artificial Intelligence. Freefall" — third edition, free download. If you still have questions, reach out to us for training or consulting.


