Before the use case: Is your data ready?
The order you work in decides more than the AI model you pick.
The wrong first question
Most AI conversations start in the same place. "Where could we use this?"
It is the wrong place to start and often begins with the tool and works backwards to find a justification. That is how organisations end up with pilots that demo well and change nothing.
There is a better order, and it has four steps. Beginning with; What problem are we solving. Does this process still make sense. Do we have the data. Then, and only then, where does AI fit.
Skip a step and you do not go faster. You get an expensive answer to a question nobody asked.
Step one: what problem are we solving
Start with a decision, not a department.
The useful question is not where our inefficiencies are. It is which decisions we make badly, slowly, or without enough information, and what better would look like.
Answer three things in writing before anything else happens:
- What decision are we trying to improve, and who owns it today? If no name is attached, you are not ready to automate it.
- How would we know it improved? Pick a measure someone already cares about.
- What is it costing us now? Rework, delays, escalations, risk carried. If nobody can size it, support will fade the first time the project gets hard.
This step is skipped almost everywhere. It is also the cheapest way to kill a bad idea.
Step two: does this process still make sense
Automating a broken process gives you a faster broken process, and now it is harder to see because it happens inside a system.
So before asking what a machine could do here, ask what should be happening here at all.
- If we designed this today, would we build it this way? Many steps exist because of a constraint that no longer applies.
- Which steps create value, and which fix an earlier mistake? Checking, chasing and re-entering data are usually symptoms, not work.
- What could we stop doing? Sometimes the best answer needs no technology at all.
There is a second reason to do this now. A redesigned process produces different data from the one you have today. Decide what you need to capture while you are designing it, not two years later when you find it was never recorded.
Step three: do we have the data
This is the step that decides everything after it, and it is usually treated as a technical detail. It is not. It belongs to whoever owns the decision from step one.
-
What data does this need, and how much? Start from the decision and work outwards. Not from the data you happen to hold. Volume matters less than coverage: does the data include the unusual conditions, or only normal operation? A system trained on good days will fail on a bad one.
-
How is it collected, and by whom? Data carries the conditions it was collected in. Manual entry brings typing errors. Sensors drift. A field filled in at the end of a long shift is not the same as one filled in carefully. If a person can influence a number, know what pressure they are under when they enter it.
-
Where does it live, and who can get to it? This is the most common problem. Not missing data, but unreachable data. Sitting in a system nobody can query, owned by a team nobody asked, in a format that needs a specialist to open.
-
Does it carry its context? A number on its own is not information. Batch, time, equipment, operator, conditions, units, version. Context is what lets you compare one value to another, and it is the first thing lost when data moves between systems.
-
Can you trace where it came from? If you cannot follow an output back through the data to its source, you cannot explain it, correct it, or defend it. In GxP settings this is a formal requirement. Everywhere else it is the difference between a system you can trust and one you take on faith.
-
What rules follow it across borders? Where data is created, processed and stored can each carry different obligations. For organisations operating globally this is routine, not an edge case. Plan for it early, because moving data later is painful.
-
How much of it is redundant, obsolete or trivial? Most organisations store a great deal that is duplicated, out of date, or was never useful. It costs money, widens your risk, and weakens anything that reads across it.
-
Who owns it? Every important dataset needs one named person responsible for its quality, its access rules and its lifecycle. Not a committee. This is the most common gap and the easiest to fix.
Step four: now add AI
Once the first three steps are answered, the AI question becomes much narrower than people expect. The instinct is to ask which model is best. The better question is what kind of system this decision can tolerate.
Draft Annex 22, the EU's first GMP guidance written specifically for AI in medicines manufacturing, makes this point clearly. As drafted, it expects static models in critical applications: parameters fixed, behaviour repeatable, outputs explainable. It indicates that adaptive and probabilistic models, including generative AI, should not be used there at all. Non-critical use cases have scope to consider probabilistic models, with qualified people accountable for whether the output is fit for purpose.
The text is not final. Consultation closed in October 2025, the EMA held an expert workshop in mid-2026 on whether risk-based safeguards could make room for probabilistic models, and the drafting group is still considering it.
But the underlying logic will hold whatever the final text says. The state of your data and your model does not just affect how well a use case performs. It decides what the use case is allowed to be.
That changes data readiness from housekeeping into scoping. It is not a tidying job that delays the interesting work. It decides which options are available at all. If you cannot show what the data is, where it came from, how it is controlled and why the output is repeatable, the question is no longer which use case is better. That use case is off the table.
Most organisations have no Annex 22 to work from. They have something harder: a critical line nobody has drawn. So draw it. Identify the decisions you could not defend if the output was wrong. Above that line you need repeatability, traceability and a named human owner. Below it, you have more freedom than you are currently using.
Then ask three questions about the AI itself:
- Are we informing a decision, recommending one, or making one? Three very different risk profiles, often described in the same words.
- Who checks the output, when, and against what? If the answer is that someone will probably notice, you do not have oversight.
- What happens when it is wrong? Not if. A system with no defined failure path is one that fails quietly.
The order is the strategy
None of this is an argument for going slower. It is an argument for starting in the right place, which is usually faster and always cheaper.
Problem. Process. Data. Then AI.
Organisations that work in that order end up with fewer AI projects and better ones. They can say why each exists, what it improved and who owns it. Organisations that start with the tool end up with a pile of pilots and a vague sense that AI is not working for them.
The difference is rarely budget or talent. It is the order.
Three questions for your next planning conversation
- Which decisions could you not defend if the output was wrong, and who owns each one today?
- For the use case you most want to build, could you show where the data came from, how it is controlled and why the output is repeatable?
- If the answer is no, is that a reason to fix the data, or a reason to pick a different use case?
If those questions are live in your organisation, we would be glad to talk. The Institute of Applied AI works with life sciences organisations on this sequence, from data readiness through to governance that holds up under inspection.
About the author: Siobhán O'Leary is an Applied AI Advisor and co-founder of The Institute of Applied AI, helping organisations build AI capability that is grounded in literacy, governance, and practical adoption. She holds the AIGP certification from the IAPP and publishes AI in Motion, a weekly newsletter for leaders navigating AI beyond the headlines.
