Unready Data Means Unusable AI: A Data Readiness Framework for Government

Short answer: Data readiness is the degree to which data is ready for a specific use, such as analysis, decision-making or training an AI model. Thai government agencies hold large amounts of data, but problems such as silos, duplicate records, missing fields, legacy databases and biased samples keep it from being usable. NXT's framework assesses readiness across five dimensions: governance, quality, architecture and integration, security and privacy, and data literacy and culture. Agencies should invest in data readiness before investing in AI.

Why do so many AI projects never make it into real use? The answer is usually not technology but data

Whenever AI comes up, government agencies tend to think of cutting-edge algorithms, expensive platforms or big-name vendors. But the truth many organizations have learned the hard way is this: however good the AI is, it can only be as good as the data fed into it.

The Thailand AI Readiness Assessment Report 2025 by UNESCO and TDRI states that Thailand has more than 28,000 open government datasets, but their quality and reusability remain limited. Having a lot of data does not mean the data is "ready to use."

The Big Data Institute (BDI), a public organization under the Ministry of Digital Economy and Society, likewise confirms that the main problem in the Thai public sector is not a lack of data, but that the data is not in a state where it can be put to use.

This article walks you through a Data Readiness Framework that government agencies can use to assess and raise their data readiness, paving the way for effective use of AI.

What is data readiness, and why does it matter more than data collection?

Data readiness is the degree to which data is ready to be used for a specific purpose, whether that is analysis, decision-making or training an AI model.

Research published on PMC (National Library of Medicine) proposes a data readiness framework covering four main dimensions:

  • —Quality: data is accurate, complete, consistent and reliable
  • —Availability: data can be accessed when it is needed
  • —Interoperability: data from different sources can be linked together
  • —Provenance: the origin and reliability of the data can be verified

Government agencies that have been collecting data for 10–20 years may hold enormous amounts of it, but if that data has quality problems, it amounts to "having resources you cannot use."

The 5 most common data problems in Thai government agencies

From our experience working with public-sector agencies, and from BDI data, we see the same problems recur again and again:

1. Data silos: data scattered across departments

Each department, division or bureau has its own database, uses different formats and field names, and there is no central system linking them. As a result, when data needs to be analyzed across units, more time goes into "cleaning the data" than "analyzing the data."

2. Duplicate records

The same citizen may appear in several databases with different spellings of their name, outdated addresses or mismatched reference numbers. AI trained on duplicated data will produce skewed results.

3. Missing fields: incomplete data

Data collection forms change over time. A field that used to be optional becomes important for AI, but years of historical records don't have it. So a dataset that looks large actually contains very few usable records.

4. Outdated databases: legacy systems that won't die

Many agencies still run database systems developed 15–20 years ago. Migrating the data is costly and risky, so it keeps getting postponed until it becomes accumulated technical debt.

5. Biased samples: data that isn't representative

Some datasets over-represent certain areas or population groups. AI trained on biased data produces unfair results — and for public services that must serve every group of citizens equally, that is an unacceptable risk.

A data readiness framework for Thai government agencies

We propose a Data Readiness Framework designed specifically for the Thai public-sector context, synthesized from several sources: BDI's Government Big Data Analytics Framework, UNESCO RAM, the World Bank Open Data Readiness Assessment (ODRA) and international best practice:

Dimension 1: Data governance

Data governance maturity levels
LevelStatusCharacteristics
Level 1: InitialNo policyNo one owns data; no data management policy
Level 2: DevelopingPartial policySome departments have a data owner; a policy exists but is incomplete
Level 3: DefinedClear policyA complete data governance framework; a data steward in every department
Level 4: ManagedMeasurableData quality KPIs; regular monitoring
Level 5: OptimizedContinuous improvementAutomated data quality checks and a continuous improvement process

Where government agencies should start: Appoint a Chief Data Officer (CDO) or a clearly designated person responsible for data, build a central data catalog, and write a data governance policy aligned with the PDPA (Personal Data Protection Act).

Dimension 2: Data quality

Data quality criteria
CriterionDescriptionExample problem in government
AccuracyData matches realityCitizens' addresses don't match where they live now
CompletenessData is complete, with no blank fieldsMore than 40% of phone numbers are blank
ConsistencyThe same data matches across every systemA name in database A is spelled differently in database B
TimelinessData is currentIncome data is from last year
UniquenessNo duplicate dataOne citizen has 3 records

Measurable target: Set a data quality score for each important dataset. A score above 80% is considered ready for AI; below 50% needs improvement before use.

Dimension 3: Data architecture and integration

Good data architecture is the foundation of AI that actually works. Its main components are:

  • —Data catalog: a central index of data showing what data is where, who owns it and what quality level it is at
  • —Master data management (MDM): a system for managing "master data" such as citizen and organization data so it follows one standard across the agency
  • —Data integration layer: a linking layer that lets data from multiple systems "understand each other"
  • —API gateway: a standard channel for exchanging data between agencies

BDI's B.I.G (Big Data Integration and Governance) project, which links data across agencies through platforms such as Health Link, Travel Link and Envi Link, is a good example of data integration at the national level.

Dimension 4: Data security and privacy

In a public-sector context that handles large volumes of personal data, data security cannot be neglected:

  • —PDPA compliance: every dataset must comply with the Personal Data Protection Act
  • —Access control: set data access rights according to roles and duties (role-based access control)
  • —Data anonymization: using data to train AI requires an appropriate de-identification process
  • —Audit trail: it must be possible to check who accessed what data, and when

Dimension 5: Data literacy and culture

However good the technology and policies are, they mean nothing if people in the organization don't value data:

  • —Staff data literacy: people at every level must understand why good-quality data matters
  • —Data-driven decision-making: decisions must be based on data, not just experience or "gut feel"
  • —Accountability: everyone knows they are responsible for maintaining data quality
  • —Continuous learning: there are ongoing programs to build data skills

Data readiness assessment checklist

Use this checklist to assess your agency's data readiness:

Data readiness assessment checklist
No.ItemYesNo
1We have a written data governance policy☐☐
2Important datasets have a clearly designated data owner/steward☐☐
3We have a data catalog that lists all of the organization's datasets☐☐
4Data quality is measured regularly☐☐
5Data from different departments can be linked together☐☐
6There is a central data format standard that every department shares☐☐
7Data complies with the PDPA☐☐
8There is a role-based access control system☐☐
9Staff have basic data literacy☐☐
10There is a systematic data cleansing process☐☐
11Data lineage can be tracked (we know where data comes from)☐☐
12There is a plan to build staff data skills☐☐

How to read the results:

  • —10–12 "yes" answers: high data readiness — ready for an AI project
  • —6–9 "yes" answers: a foundation is in place, but there are important gaps to close before starting on AI
  • —Fewer than 6 "yes" answers: low data readiness — invest in data management before investing in AI

Roadmap: from unready data to AI-ready data in 4 phases

Phase 1: Data discovery and assessment (months 1–2)

Survey and assess the state of all the organization's data, build an initial data catalog, identify the most important datasets (priority datasets) and assess the data quality of each.

Phase 2: Data foundation (months 3–6)

Establish a data governance framework, set data standards, appoint data owners/stewards, start data cleansing for priority datasets and create master data for core data.

Phase 3: Data integration and quality (months 7–12)

Link data across departments with an integration platform, implement automated data quality checks, build data pipelines that support AI use and run a data literacy program for staff.

Phase 4: AI-ready data operations (month 13 onward)

Build a data-as-a-service layer that gives AI teams immediate access to good-quality data, implement continuous data quality monitoring and create a feedback loop from AI results back to improving the source data.

Quick wins: 3 things you can do now without waiting for extra budget

You don't need to wait for a big plan. Start with these three:

1. Build a data inventory: Survey what datasets the organization has, where they are and who is responsible for them. Even a simple spreadsheet is better than nothing.

2. Pick one important dataset to clean: Choose the dataset that is used most often or matters most, run data profiling to see what problems it has, then start fixing them systematically.

3. Appoint a data champion in each department: This doesn't have to be a data scientist, but someone who understands their department's data and can act as the go-between in pushing data quality forward.

"Investing in AI before sorting out your data" = money wasted

If the data isn't ready, an AI project will fail no matter how expensive it is. This isn't theory; it's a lesson many agencies around the world have already learned.

Data readiness is the foundation you must invest in first. It isn't something you do "alongside" AI, but something you must do "before" investing in AI.

NXT Consulting Group helps government agencies assess data readiness, put a data governance framework in place and build a roadmap from scattered data to data that is ready for AI. Our work is assessment and planning; we do not build data systems.

Talk to us about assessing your organization's data readiness →

Translated from the Thai original by NXT Consulting Group.

Share this article