Data Preparation Before Any AI Project Starts
The data decides the outcome
Projects of this kind rarely fail because the model was unsuitable. They fail because nobody knew what the organisation’s data contained until somebody tried to train, retrieve or automate with it — and what it contained turned out to differ from what everybody assumed.

Preparation is where those assumptions get tested. It is also the stage most likely to be compressed, because it produces no visible output and a demonstration is already booked. In our experience at SmartEdge IT Solutions, the discovery note that comes out of this stage is more useful to a business than the working system that follows it, because it answers questions about operations that were previously unanswerable.
None of it requires sophisticated tooling. Most of it is a person reading records carefully, which is unglamorous, slow, and consistently the part that determines whether the rest works.
Inventory before opinion
The first task is not analysis, it is enumeration. List every system that holds relevant information, including the ones nobody thinks of as systems.
For a small business that typically means the accounts package, the CRM, the shared inbox, the shared drive, the website content, the spreadsheets that have quietly become load-bearing, and the person who carries institutional knowledge but has no system at all. That last category deserves particular attention, because it is often the only source for something the business relies on heavily.
For each one, record five things in a table: what it holds, roughly how many records, when it was last updated properly, who can read it, and whether anybody would object to that data being processed by a third party. The last column frequently determines the scope of a project, and it is rarely the one anybody thinks to ask about.
Then look at the overlap. Two systems holding the same customers with different identifiers. One spreadsheet a month out of date. An export that runs every Sunday and breaks silently when somebody adds a column. These findings shape the project, and they come out of the inventory rather than out of the tooling discussion.
This stage is where a lot of the perceived scope disappears. Questions that had been scheduled as features turn out to be reporting problems, and reporting problems turn out to be data problems. Better to learn that in week one. SmartEdge IT Solutions will usually spend longer on this inventory than a client expects, and it is the part that determines the size of the quote.
Quality and suitability are different things
Data quality gets discussed as a list of defects — missing fields, inconsistent formats, duplicates, stale records — and all of those are real and all of them are fixable to some degree. But the defects that block a project are usually about suitability rather than tidiness.

A customer support dataset can be immaculate and still unusable for a retrieval system, because every answer depends on the state of an order at the time it was written. Free-text notes with no structure are fine for search and useless for anything that needs to filter. An export from two years ago may be beautifully clean and describe a product line that no longer exists.
So the question to ask is not “is this data clean” but “what is this data a record of, and is that still true”. The reframing tends to change the plan, because it moves the conversation from cleaning to provenance.
Two questions do most of the work here: what decision or action will this data support, and how much of it actually relates to that action. The second one is usually the surprise. Quite often the answer is that under half the records concern the case in hand and the remainder is noise that adds cost and reduces accuracy.
Where the relevant proportion is small, it is often more honest to narrow the project than to clean the remainder. If the interesting cases are a fifth of the volume, a workflow scoped to that fifth may serve the business better than one that also handles the other four.
Build the evaluation set first
Before any preparation, assemble a set of cases with known correct answers. Twenty is often enough to start; a hundred gives you something you can actually track over time.

This is the step teams skip, because it feels like duplicated work. It is not. Without a fixed set of questions and answers there is no way to tell whether a change improved anything, and every subsequent improvement becomes a matter of opinion. With it, you can compare two retrieval strategies, two models or two chunking settings and get an answer rather than a preference.
Take the cases from real history where possible: the queries people actually typed, the documents that were actually approved, the situations that actually escalated. Include the awkward ones deliberately. A set of straightforward cases will tell you the system works, which is the least interesting result available.
Keep the set versioned and separate from the working data. It has to survive pipeline changes, because its entire purpose is to be the fixed point that everything else moves against.
On small projects this doubles as an acceptance test the business can read, which is worth more than a technical metric when you are agreeing scope with an operations lead.
Labels and ground truth
Any project that trains, tunes or evaluates against human judgement needs labels, and labels are where the hidden cost sits. They need people who know the work, they need a written definition, and the definition will be argued about at greater length than the labels themselves.
The definition comes first and takes longer than expected. If three competent reviewers would disagree on whether a ticket is a billing issue or a technical one, the labelling exercise will produce noise and everybody involved will be frustrated. Write the categories, write the boundary cases, and accept that some real cases will not fit cleanly into any of them.
Then label a small batch with two people independently and compare. Disagreements tell you precisely what the definition is missing. Fixing a definition after a hundred labels is expensive; fixing it after ten is free.
Be honest about volume. If a useful model needs thousands of consistent labels and the budget covers eighty, the correct decision is to change approach rather than accept a badly labelled set. Retrieval over well-organised material, with weak supervision and a small amount of human labelling, frequently beats a small supervised model trained on inconsistent labels — a route worth considering before committing to a labelling programme at all.
Things that look fine until you look
A handful of checks catch most of what goes wrong later, and none of them needs specialist tooling.

- Duplicates that are not identical. The same document filed twice under slightly different filenames, or the same customer created independently by two integrations. Exact-match deduplication finds none of these.
- Near-identical boilerplate. Long shared headers and footers that dominate similarity scores and pollute any retrieval based on them.
- Records where the label is a timestamp. Common where systems export status history; sorting by date is a poor proxy for the label you actually want.
- Silent staleness. An export that runs on a schedule and fails quietly, leaving a three-week-old snapshot that looks complete.
- Encoding damage. Text extracted from PDFs and images with mangled characters and broken word boundaries, which then quietly poisons anything searching over it.
- Language and register mixing. Material in several languages or formats handled as one corpus, producing confident answers in the wrong language.
Look at actual samples rather than counts. Ten records printed in full tell you more than a table of field-completion rates, because the defects that matter are usually in the text and in the relationships between records, not in the columns.
Access, ownership and who may see what
Before data moves anywhere, four things need answering in writing, and these are commercial questions rather than technical ones.

Who owns the data and may it be processed by a third party for this purpose? What contractual basis applies, and does it differ between data about customers and data about staff? How long may it be retained, including inside logs and backups, and how is deletion verified rather than merely requested? And which categories must never leave the environment, whether because of health information, payment data, or a specific obligation the organisation has accepted?
Practically, this shapes architecture more than anything else. If answers must not leave the organisation’s own systems, retrieval has to run on premises or under an arrangement that satisfies the requirement, which changes cost and effort considerably. Better to establish that in the first week than in the third month.
Also worth settling: whether derived outputs become an asset. If a system produces classifications or summaries about customers, is that information retained, and does the same restriction apply to it? Treat it as the same category as the source unless there is a specific reason not to.
Our system integration work on client estates usually surfaces these questions early, because the access model of existing systems constrains what any automation can reasonably reach.
When the answer is that the data is not ready
Sometimes the right output from a discovery exercise is that a particular project should not start yet, and that is a perfectly good outcome. The alternative is spending the budget learning the same thing more expensively.
The signals are recognisable. The source data does not exist and would have to be created. The field that matters most is maintained by hand in one spreadsheet by one person who is leaving. The definition of success depends on a judgement nobody has written down. The volume is too small for the approach to be worthwhile, or the variation is so wide that any model would be guessing.
Where that is the conclusion, say what would make it ready and roughly what that would take. Often it is unglamorous: capture a field consistently for a quarter, agree a definition with the operations team, fix an integration that has been failing silently, obtain a consent nobody asked for. That work has value in its own right even if the project never happens, because it makes a process visible for the first time.
There is also a version of this that is less about data and more about scope: the project is ready, but smaller. Most small business automation efforts that succeed did less than the original proposal on the first pass. Cutting scope until the data supports it is a design decision rather than a failure. And where the fix is a clearer process rather than cleaner data, business process automation scoping is usually the more useful conversation to have.
What ends up in the discovery note
The deliverable of this stage is a written document, and it is worth resisting the urge to turn it into a presentation. It should be readable by the operations lead without translation.

It records the systems inventory and what each one is actually good for. The defects found and what was done about them. The proportion of records relevant to the intended use. The evaluation set, its size and where it lives. The access and ownership position, including anything still unresolved. The realistic scope, written as tasks rather than as themes. And the explicit list of what is out of scope.
That last item does more work than any other. The second idea will arrive within a month and it will sound perfectly reasonable, and a written boundary agreed before anybody is emotionally invested is the only thing that makes the answer “no” easy to give.
If the work needs more than a couple of people can do alongside their day jobs, it becomes a project rather than an exercise, and our AI automation engagements tend to begin from this note rather than from a tool demonstration. Where the heavy lifting is reshaping and moving data rather than reasoning about it, workflow and data automation is the more accurate description.
Read the note again six months after go-live. It is usually the most reliable record of what the business assumed at the start, and comparing it against what happened is the most useful evaluation of the whole exercise available.
