Almost every company that asks us for a data strategy has already bought the tool. They have BI licences, a half-filled cloud warehouse and three dashboards nobody has opened since the demo. What's missing isn't software: it's an agreement on which number is the right one, who answers for it, and in what order the mess gets cleaned up. That is a data strategy, and it's a business document, not an architecture diagram.
This guide covers how that agreement gets built in an ordinary mid-sized company, with an ageing ERP, a half-finished CRM and plenty of spreadsheets. What a data strategy actually is and isn't, why the ones that start with technology fail, how to settle the source of truth, how to decide what to fix first, what a realistic 90-day plan looks like and what to measure to know it's working.
What is a data strategy and what isn't it?
A data strategy is the written answer to four questions: which decisions do we want to make better, what data do we need for them, what state is that data in today, and in what order will we fix it. It fits in five or six pages. If it runs to forty, it's usually a wish list in disguise.
What it is not:
- Not an architecture diagram. The boxes-and-arrows drawing is a consequence, not a starting point. Choosing the warehouse before knowing which questions need answering is picking the road before the destination.
- Not an inventory of every system you own. An exhaustive inventory takes months and ages in weeks. Mapping the systems that touch the three or four decisions that matter is enough.
- Not a three-year plan. Nobody knows which tools will exist in three years. What works is a one-year horizon with the first quarter detailed day by day.
- Not an AI project. AI is just one more data consumer — a demanding, noisy one. If the strategy doesn't hold up without AI, it won't hold up with it either. That's the argument we develop in data before AI.
The acid test is simple: put the document in front of your CFO. If they can't point to three concrete decisions that will involve less arguing six months from now, it isn't a strategy yet. It's an infrastructure budget looking for justification.
A data strategy isn't judged by the elegance of its architecture, but by how many meetings stop opening with "that's not the number I have."
Why do data strategies that start with the tool fail?
Because they solve the storage problem and leave the agreement problem untouched. We've seen the same pattern often enough to describe it in stages.
First comes the platform. Someone stands up the warehouse, connects the ERP and publishes a sales dashboard. A month later, sales says the revenue figure doesn't match theirs. You investigate: sales counts the order when it's signed, finance when it's invoiced, operations when it's delivered. None of them is wrong. Nobody ever decided which of the three definitions is the official one.
From there the ending is predictable. Each department goes back to its own spreadsheet "because mine adds up", the dashboard becomes decoration, and the project gets filed as a technology failure when it was really a governance failure. The cost isn't just the licence: it's that the next time someone proposes cleaning up the data, management already has a bad memory attached to the idea.
The four mistakes we see repeated most:
- Starting with the hardest data. Product margin with indirect costs allocated is usually the worst possible first project. Enormous effort, arguable result, months with nothing to show.
- Modelling everything before delivering anything. Six months of a perfect data model with not a single report in production guarantees the sponsor loses interest before the launch.
- Not naming owners. Data with no owner doesn't get corrected — it gets commented on. And comments don't change the source system.
- Confusing a report with a decision. Publishing the number changes nothing if nobody has the authority or the habit of acting on it. It's the same mechanism behind why AI pilots fail.
How do you decide what the source of truth is?
The source of truth isn't a system, it's a decision per relevant data point: for data X, system Y wins, with this definition and this owner. It's documented in a boring table that's worth more than any diagram.
| Data | System of record | Agreed definition | Owner | Frequency | |---|---|---|---|---| | Customer | CRM | Entity with a unique tax ID; sites are locations, not customers | Sales director | Daily | | Order | ERP | Confirmed when a firm order number exists | Operations | Daily | | Revenue | ERP (billing) | Invoice issue date, excluding VAT | Finance | Daily | | Product | Item master | Active item code; variants hang off the parent | Product | Weekly | | Unit cost | ERP + costing | Standard cost for the period, not actual per batch | Financial control | Monthly |
Three practical rules for filling this in without getting stuck.
One discussion per data point, with a decision-maker in the room. The classic mistake is opening an email thread to agree what a customer is. Book forty minutes, bring the person who can decide, and leave with the definition written down even if not everyone likes it. An imperfect published definition beats a perfect pending one.
Document what the definition leaves out. "Revenue does not include orders pending delivery" heads off half the future complaints. The explicit boundaries are the real deliverable.
Accept living with two figures when both are legitimate. Signed sales and invoiced sales are both useful. What you can't do is call them the same thing. Different names, different definitions, and both on the same panel so nobody gets suspicious.
Once this table exists and is maintained, you're already doing data governance without having set up a committee. And if you do need one, at least it will have something concrete to decide on.
How do you prioritise which data to fix first?
Along two axes: the value of the decision it unlocks and the real effort of cleaning it up. Rank the candidates and go after the high-value, medium-or-low-effort ones. It sounds obvious; almost nobody does it, because the conversation tends to drift towards which data is dirtiest, which is a different question.
Questions that help score each candidate:
- Which concrete decision changes? If the answer is "we'd have visibility", it scores nothing. If it's "we'd stop restocking this product family by gut feel", it does.
- How often is that decision taken? Data feeding a weekly decision is worth far more than data for the annual board meeting.
- How many people recalculate it by hand today? That number is your most defensible saving and the easiest to measure.
- Does the source data exist, or must it be created? Cleaning is expensive; capturing something nobody records today is far more expensive and disrupts operations.
- Who will push back? If the owner of the source system doesn't want it touched, add that to the effort. Organisational resistance is a cost line, not a footnote.
A typical order that comes out of this exercise in a distribution business: customers and orders first, because almost everything hangs off them; then stock, because it unlocks weekly replenishment decisions; then margin, which is what management asked for on day one and which depends most on the other two. Explaining that order with this logic stops it looking like the data team doing whatever it fancies.
If your medium-term ambition includes forecasting or predictive analytics, add one more criterion: available history. A model needs two or three years of data under the same definition. If you changed ERP eight months ago, that use case drops down the list however exciting it sounds.
What does a 90-day data plan that holds up look like?
Ninety days is enough to resolve one decision end to end. Not three. This is how we split it and where the effort actually goes.
| Phase | What happens | Tangible result | |---|---|---| | Weeks 1-2 | Business interviews; pick 1 decision and 2 metrics | One-page document with owner and baseline | | Weeks 3-4 | Source-of-truth table for the data involved | Definitions signed off by their owners | | Weeks 4-7 | Ingestion and minimum model; only what feeds those metrics | Data refreshed daily, no manual steps | | Weeks 7-10 | Panel with the 2 metrics and their breakdowns; business validation | Figures that match known reality | | Weeks 10-12 | Usage routine: who looks at it, when, and what they do | The decision is made with the panel open | | Week 13 | Review: what to scale, fix or drop | Next quarter prioritised on evidence |
Two warnings about this calendar. First: weeks 10-12 are the ones most often skipped and the only ones that are non-negotiable. A correct panel that never made it into a recurring meeting dies on its own, and it dies labelled "the data project didn't work".
Second: leave slack for whatever turns up. Ingestion always surfaces something — a field used for two purposes, duplicate customers under different tax IDs, dates in three formats. If the plan is booked at 100 % capacity, the first discovery derails it. Planning to 70 % is more honest and finishes sooner.
On spend: in this first quarter, infrastructure is usually the smallest line. The dominant cost is your own people's hours spent defining, validating and adopting. Budgeting only for licences and technical hours is the most common way for a project to run out of fuel in week eight.
What roles does a data strategy actually need?
Fewer than the textbook says, but none of them can sit empty. In a mid-sized company these are hats, not full-time headcount.
- Sponsor. Someone in senior management who answers for the outcome and unblocks cross-department arguments. Without this role, any disagreement lasts weeks.
- Data owner. For each critical data point, the business person who decides its definition and answers for its quality. Not IT: the person who knows the operation.
- Technical data profile. Whoever builds the ingestion, the model and the validations. Internal or external, but stable: rotating here erases accumulated context.
- Analyst or translator. Whoever turns the business question into a metric and explains the result without jargon. It's the most undervalued role and the one that most determines adoption.
With those four hats and a thirty-minute monthly review you have enough governance to start. Committees, catalogues and formal policies come later, when there's something to govern. Building the bureaucracy before the first clean data point is the fastest way to exhaust the organisation's patience.
How do you measure whether the data strategy is working?
With indicators that don't depend on anyone's opinion. Four cover most situations:
- Time to the number. How long it takes from someone asking to a reliable figure existing. Going from three days to one hour means the strategy is working.
- Open discrepancies. How many times a month two departments present different figures for the same concept. It should fall towards zero for the data already agreed.
- Manual recalculation hours. The ones your people spend rebuilding reports in spreadsheets. It's the most tangible saving and the easiest to defend upwards.
- Actual usage. Distinct users opening the panel in the week the decision gets made. Without usage, everything else is set dressing.
Set the baseline before you start, even roughly and by survey. Without that starting point, any improvement will look arguable and you won't be able to justify the next quarter. Once these numbers exist, the conversation about what business intelligence is and which tool to buy becomes easy, because you already know what it has to solve.
Frequently asked questions
How long does a data strategy take to produce results?
The document itself takes two or three weeks including interviews. The first tangible result — a decision taken with reliable data instead of a spreadsheet — fits in a quarter if you scope it to a single decision. Plans that promise to transform all reporting in three months don't end well.
Do I need a data warehouse to have a data strategy?
Not to have one; often, yes, to execute it beyond the first use case. Many companies get through their first quarter with tidy extracts and a simple model. The warehouse earns its place when several sources need cross-referencing daily, or when you need history the source system doesn't keep.
Who should lead the data strategy, business or IT?
Business sets the decisions and the definitions; IT answers for the data arriving clean and on time. If IT leads alone, you get correct systems nobody uses; if business leads alone, you get demands with no foundation. The sponsor belongs in senior management, not in the systems department.
Is a data strategy worth it for a small company?
Yes, and it takes up less space. In a twenty-person company it can be a six-row source-of-truth table and two agreed metrics. The value isn't in the size of the document, but in there being one single definition for the three or four numbers the business is run on.
If you're about to buy a data platform and haven't written down which decisions it needs to improve, that's exactly the gap the audit fills: looking at your systems, your definitions and your meetings to come out with a source-of-truth table, one prioritised use case and a 90-day plan you can defend to the board. If you'd rather test the idea in half an hour first, let's talk.
Shall we apply it to your case?
The 360° AI Audit turns these ideas into a concrete plan for your company: three weeks, fixed price and the full picture of your AI before spending a euro.
See the 360° Audit→ Let's talk↗