A digital twin is a replica of a physical asset (a machine, a line, a plant) fed with real data from that asset, used to answer questions without touching the asset. That's the useful definition. Everything else — the 3D model spinning on a screen in the meeting room — is an expensive animation. The difference isn't the software: it's whether there's a live data flow behind it, and whether anyone makes different decisions because of what the model says.
This guide is for plant managers, maintenance leads, engineering firms and mid-sized manufacturers who have just been pitched a digital twin. It covers the levels that exist, the data each one demands, when it genuinely pays off, what it costs in orders of magnitude, and how to run a first case on a single machine instead of starting with the whole factory.
What is a digital twin, and what isn't one?
Three pieces have to be present before you can call it a digital twin:
- A model of the asset: its geometry, its physical behaviour, its process rules, or some combination. It can be a CAD model, a discrete-event simulation or a statistical model trained on history.
- A data flow from the real asset: sensors, PLC, SCADA, MES, production counters, maintenance records. Frequent enough that the model doesn't lag behind reality.
- A business question the model answers: will this bearing hold until the August shutdown? Where's the bottleneck if I add one more product reference? What do I lose if I drop line speed by 8%?
If the data flow is missing, it's a simulation model — useful, but not a twin. If the question is missing, it's a visualisation project. And if the model is missing, what you have is a plant dashboard, which is usually the sensible and far cheaper first step: we cover that in executive dashboards.
A digital twin isn't a pretty screen of the factory. It's a model that commits: it predicts something checkable and gets it wrong in measurable ways.
What it doesn't do: it doesn't replace the maintenance lead, it doesn't tune the machine on its own (except in very mature deployments with a deliberately designed closed loop), and it doesn't fix a process nobody has standardised. A twin of a chaotic process reproduces the chaos with more decimal places.
What levels of digital twin exist?
The term covers things that differ in cost by two orders of magnitude. Ask the vendor to state which level their proposal sits at:
| Level | What it does | Data required | Relative cost | |---|---|---|---| | 1. Descriptive replica | Shows the asset and its current state (stoppages, speed, temperature) | PLC/SCADA signals in near real time | Low | | 2. Connected model | Simulates behaviour and compares it with reality: flags deviations | The above + a validated process model | Medium | | 3. Predictive model | Anticipates failures, quality issues or bottlenecks from history and physics | The above + 12-24 months of history with failure labels | High | | 4. Prescriptive twin | Proposes (or applies) a setpoint change and evaluates scenarios | The above + write integration with control and risk governance | Very high |
Most mid-sized companies capture almost all the return at levels 1 and 2. Level 3 pays off when failures are expensive and recurrent. Level 4 is a minority, and demands a maturity in automation and security you can't improvise.
A common mistake is buying level 4 in the contract and ending up at level 1 in practice, because the data never arrived at the frequency required. Always ask what happens if data arrives every 15 minutes instead of every second: if the answer is "then it doesn't work", you've found the project's risk.
What data and sensors does it actually require?
This is where half of these projects collapse. A twin demands three things many plants don't have in order:
- Continuous signal from the asset. A piece counter isn't enough. A wear model needs process variables (vibration, temperature, current draw, torque, pressure) sampled at a rate consistent with the phenomenon you want to see. A bearing degrading over weeks doesn't need 10 kHz sampling; a vibration fault does.
- Labelled event history. Knowing when it failed, what was replaced and why. If maintenance records are free text in a spreadsheet, there's upfront normalisation work that isn't optional. No labels, no predictive model — only generic anomaly detection.
- A coherent asset master. The same machine named identically in the CMMS, the MES and the ERP. It sounds trivial and it's the number one reason the numbers don't add up. We see it in almost every project: see data quality.
Then add the infrastructure side: OT network separated from IT, a gateway to pull data off the PLC without compromising control, storage with enough retention (two years of high-frequency history is bulky, and that costs money every month). That plumbing is pure data engineering and usually accounts for 50-70% of the real effort.
Rule of thumb: if you can't export, within 48 hours, a CSV with one year of the five key variables from the candidate machine, you're not ready for a level 3 twin.
When is a digital twin premature spend?
Clear signs you should wait and do something else first:
- No failure history. If the machine has been installed for eight months and failed twice, there's nothing to learn. Instrument it properly and collect data; the model can come in a year.
- The process changes every month. If every order reconfigures the line, the model is obsolete before it's validated. Standardise first, model second.
- The real problem is organisational. If the bottleneck is that procurement doesn't flag delays, a twin won't fix it. A twin of a coordination problem is an expensive mirror.
- Nobody will look at the output. Without a plant owner who uses the prediction to change the shutdown plan, the project dies within six weeks, like any tool without an owner: we explain that in change management.
- The asset is cheap to replace. Modelling a €1,200 pump that gets swapped in two hours doesn't pay, however technically feasible it is.
The honest alternative in those cases is usually humbler and more profitable: add sensors, build an OEE and downtime dashboard, clean the asset master, and revisit the conversation in six months with data on the table.
Where does a digital twin pay off in a plant?
Four families of use case cover almost everything that makes economic sense:
| Use case | What it solves | Data required | Cheaper alternative | |---|---|---|---| | Critical equipment health | Anticipates failure of the asset that stops the line | Vibration/temperature + failure history | Inspection rounds + simple thresholds | | Bottleneck and capacity | Simulates adding a shift, a reference or a machine | Cycle times, stoppages, product mix | Spreadsheet with queueing theory | | Quality and process parameters | Links setpoints to defects and cuts scrap | Per-batch parameters + QC results | Statistical analysis on history | | Commissioning and changeover | Tests configurations without stopping the line | Validated model of the line and its constraints | Trials during planned shutdowns |
The first is directly related to predictive maintenance: a twin is, at bottom, the physics-model version of that same idea. The second is the one that surprises people most often: a well-built capacity simulation stops you buying a machine you didn't need, and that shows up in the same month's accounts.
One note on scale: pick one asset — the one that hurts most — and narrow the scope to a single question. "Can I know ten days ahead that this motor will fail?" is a project. "Let's build the digital twin of the plant" is a budget with no end.
What does it cost, and how do you run the first pilot?
No invented figures from other people's cases, but here are the components that always show up in a quote:
- Sensors: anywhere from nothing (if the PLC already measures what you need) to several thousand euros per asset if you have to add accelerometers, a gateway and cabling.
- Capture and infrastructure: OT gateway, storage and retention. That's a recurring monthly cost, not a one-off purchase, and it grows with sampling frequency.
- Model and validation: the engineering work. It depends on whether the model is statistical (weeks) or detailed physics (months).
- Integration and adoption: alerts into the CMMS, a screen on the shop floor, training the maintenance team. Always underestimated.
- Model upkeep: a model degrades when the process changes. Budget for a review at least every six months.
A 90-day sequence that works:
1. Weeks 1-2. Pick the asset and the question. Define which decision will change and what today's baseline is (how many failures, what each stoppage costs). Without a baseline there's no measurable ROI: see measuring AI ROI. 2. Weeks 3-5. Data audit: which signals exist, at what frequency, what event history there is and in what state. This is where you decide whether the project proceeds or gets postponed. 3. Weeks 6-9. Minimum model with the data available. Validate against past events the team remembers. If it doesn't catch known failures, it won't catch future ones. 4. Weeks 10-12. Shadow pilot: the model raises alerts, the team records whether it was right, nobody changes the maintenance plan yet. At the end, an informed decision to scale or stop.
The shadow pilot is the step most often skipped and the one that builds the most trust. A month of alerts that get verified convinces a maintenance lead far more than any sales demo.
Mistakes that kill a digital twin project
- Starting with the whole plant. Scope grows, time disappears into integrations, and nine months later nobody remembers the original question.
- Confusing frequency with value. Capturing everything at maximum frequency multiplies storage cost without improving the model if the phenomenon is slow.
- Leaving production out. The model is used by the people on the floor. Designed with IT alone, it's born misaligned.
- Accepting a black box. If the vendor can't explain which variables drive the prediction, the team won't trust the alert and will end up ignoring it.
- Forgetting OT security. Pulling data out of industrial control requires design: a one-way gateway where possible, network segmentation, and a review of who can write.
Frequently asked questions
How much history do I need for a predictive digital twin?
As a practical reference, 12 to 24 months including several real failures of the same type. What matters isn't the duration but the number of labelled events: two documented failures won't train anything reliable, even with three years of signal.
Does a digital twin work in a factory with old machines?
Yes, but the starting point changes. Machinery without communications needs external sensors and a gateway, which adds cost and installation work. The good news is that those are usually the assets with the most downtime, so the return shows up sooner.
What's the difference between a digital twin and predictive maintenance?
Predictive maintenance is a use case; a digital twin is one way of solving it, using a model of the asset. You can do predictive maintenance with a statistical model and no twin, and you can have a twin that does capacity simulation rather than failure prediction.
Should I buy a product or build it custom?
If your asset is standard (compressors, pumps, electric motors), there are monitoring products that cover the case with less cost and less risk. Custom development is justified when the process is proprietary and the advantage competes in the final product.
If you're weighing up a digital twin and what you genuinely don't know is whether your data can sustain one, that's exactly the question the audit answers: two weeks looking at which signals exist, what state the history is in, and which use case pays off first. And if you already know the asset and the question, let's talk about the pilot before the scope grows on its own.
Shall we apply it to your case?
The 360° AI Audit turns these ideas into a concrete plan for your company: three weeks, fixed price and the full picture of your AI before spending a euro.
See the 360° Audit→ Let's talk↗