Quick answer
Data for AI is prepared in five repeatable steps: discovery (what you have and where), cleaning (duplicates, errors, gaps), mapping (what corresponds to what across systems), transformation (a format the model can use), and validation (does the data really describe the problem). This is not a formality before the "real" project. Gartner predicts that through 2026, organizations will abandon 60% of AI projects unsupported by AI-ready data (Gartner, 2025).
Why does data quality matter more than the choice of algorithm?
Because an algorithm can only work with what it receives. A simple model trained on good data usually beats a complex model trained on noisy data, and no architecture makes up for the difference between a good dataset and a bad one.
Research points the same way from several angles. In the IDC report, data quality issues are the top reason AI projects fell short of expectations (IDC for Lenovo, 2025). In Informatica's survey, 43% of chief data officers named data (completeness, quality, readiness) as the top obstacle to moving GenAI projects from pilot to production (Informatica, 2025). RAND, based on 65 interviews with practitioners, lists lacking the data needed to train an effective model as one of the five leading root causes of AI project failure (RAND, 2024).
Meanwhile, according to a Gartner survey, 63% of organizations either do not have or are unsure whether they have the right data management practices for AI (Gartner, 2025). The problem is common, not exceptional. For other places where projects stall, see why AI pilots never reach production.
Why can't an AI model tell good data from bad?
Because it has no common sense and no business context. A model looks for patterns and correlations without asking whether they are meaningful or coincidental. It treats everything you feed it as truth.
Two examples of what that looks like:
- Inconsistent labels. A fraud detection system trained on transaction history where departments labeled fraud differently ("fraud", "suspicious", sometimes not at all). The model does not see the inconsistency. It builds contradictory rules that reduce its accuracy in production.
- Mixed units. Sales data with dollars, euros, and pounds in the same column. Without conversion to a single currency, the model either treats $100, €100, and £100 as the same number or as incomparable text. Either way it learns nothing about transaction value.
"Garbage in, garbage out" hits harder in AI than in traditional IT: the model does not just repeat errors, it generalizes them.
What are the traits of AI-ready data?
AI-ready data is data matched to a specific use case that a model can use without being misled. It has five traits:
| Trait | What it means | What happens without it |
|---|---|---|
| Clean | No duplicates, errors, or irrelevant records | A customer stored twice looks like two customers, so the model overweights traits that only "appear more often" because of duplicates |
| Consistent | The same concept is recorded the same way in every system | "Germany" in the CRM and "DE" in the ERP are two different places to the model |
| Complete | Contains the information the task needs | Without purchase history, recommendations are random; gaps in sensor data break failure prediction |
| Contextualized | Reflects the real business process | Without promotion flags, the model attributes a sales spike to customer preference |
| Governed | Has an owner, access rules, protection for sensitive data, and regulatory compliance | The team cannot legally use the data, or uses it without control |
The classic data quality dimensions (accuracy, completeness, consistency, timeliness, validity, and uniqueness) still apply. The difference is that in AI what counts is fitness for the use case: data that is fine for a quarterly report can be wrong for a model that predicts customer behavior in real time.
What does the data preparation workflow look like?
A typical workflow has five stages, and it is not linear: validation often sends the team back to cleaning or even discovery.
- Discovery. Find out what data exists, where it lives, and how it is structured. Usually customer data sits in the CRM, transactions in the ERP, and behavior in web analytics, each with its own naming conventions.
- Cleaning. Remove duplicates, fix errors, and decide how to handle gaps: drop the record, fill it with an estimate, or flag it as missing.
- Mapping. Define how fields from different systems correspond to the target structure. For example, "customer_id", "client_number", and "account_id" are the same entity.
- Transformation. Convert data into a form the model can process: scaling numeric values, encoding categories, preparing time features. In document projects this also means splitting text into chunks and adding metadata.
- Validation. Check that transformations worked, that the data really represents the business problem, and that it meets agreed quality thresholds before you train or index anything.
What about documents instead of tables?
In document-based projects such as RAG and knowledge bases, data preparation is mostly about formats. RAG (retrieval-augmented generation) is a technique where a language model first retrieves fragments of your documents and then grounds its answer in them. If the fragments are bad, the answer will be bad too.
In our agentic knowledge base, the pipeline does several things before anything reaches the search index: OCR reads Polish and English text from scans, tables are converted to text with column headers, a vision model describes charts, and every fragment gets a metadata prefix before vectorization. On top of that, Polish stemming makes "dawka" and "dawki" (dose, doses) match. In production that is over 2 million fragments from 22,786 documents. We explain the technique itself in how RAG makes AI smarter.
In other projects the same problem looks different: merging reports in different formats and matching payments to accounting entries with fuzzy matching, because the same data in two systems is never written identically.
How much does bad data cost?
More than the IT budget shows. According to Gartner, poor data quality costs organizations an average of $12.9 million a year (a 2020 estimate) (Gartner). That figure reflects large enterprises, but the mechanisms are the same at any scale:
- Operational inefficiency. People fix data instead of doing their jobs: wrong inventory levels, shipping and billing mistakes.
- Poor decisions. Forecasts and analyses built on unreliable data lead to flawed strategies.
- Eroded customer trust. Wrong addresses, duplicate accounts, inconsistent communication.
- Regulatory risk. GDPR violations can bring fines of up to €20 million or 4% of global annual turnover (GDPR, Article 83).
- Failed AI projects. No model can compensate for data that was never ready.
Checklist: is your data ready for an AI project?
Before you start talking about models, run this list for one specific use case:
| Area | Check | Ready when |
|---|---|---|
| Goal | Which decision or task will the data support? | One use case with a business metric |
| Inventory | Where is the data and who owns it? | A list of sources with owners and formats |
| Quality | How many duplicates, gaps, and errors are there? | Measured on a sample, with acceptance thresholds |
| Consistency | Are the same concepts recorded the same way across systems? | A shared glossary and ID mapping |
| Context | Does the data capture circumstances (promotions, seasonality, process changes)? | Key events are flagged in the data |
| Documents | Are scans, tables, and charts machine-readable? | OCR and extraction tested on a sample |
| Access and GDPR | Who may use the data, and on what legal basis? | Access rules, minimization, pseudonymization where possible |
| Infrastructure | Where will the data be processed? | A decision on cloud vs your own servers that meets requirements |
| Monitoring | How will you know quality is slipping? | Automated quality tests in the data pipeline |
Access rules, accountability, and compliance are the job of data governance. We cover that in a separate article: data governance, the foundation of trustworthy AI.
Can AI speed up data preparation?
Yes, and that is the good news. Language models help detect inconsistencies, normalize records, and classify documents, while vision models read scans and describe charts. That is work that used to take weeks of manual review.
There is one condition: the output of that automation has to be validated too. Rules, human-reviewed samples, and tests in the data pipeline are the only way to keep tool errors from flowing into the model.
Start with data, not the model
Companies that make AI work do not start by comparing models. They start with an honest assessment of their data: where the gaps in completeness, consistency, and accessibility are, and how to close them systematically. The quality of an AI outcome is decided before the first prompt is written or the first training epoch runs.
Before investing in a sophisticated model, invest in your data. Understand what you have, identify what you need, and do the unglamorous but essential work of cleaning, standardizing, and organizing.
Sources
- Gartner: Lack of AI-Ready Data Puts AI Projects at Risk (February 2025)
- Gartner: Data Quality, Why It Matters and How to Achieve It
- IDC for Lenovo: CIO Playbook 2025, It's Time for AI-nomics
- Informatica: CDO Insights 2025
- RAND: The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (2024)
- Regulation (EU) 2016/679 (GDPR)