
No matter how loudly a supercar with a cutting-edge engine roars, it cannot reach full speed on an unpaved road that is rough and has no signposts. It is all too easy to take a wrong turn and crash. The global race for artificial intelligence supremacy looks much the same today. Global big tech firms and governments are competing to unveil top-of-the-line supercars called large language models, but the road these cars must travel — the data infrastructure — remains poorly built. No matter how impressive the model, it is useless without a reliable network of data roads.
Indeed, the answers produced by AI, now deeply embedded in daily life and work, are not always trustworthy. Hallucinations occur frequently, with statistics of unclear origin and unclear reference dates dressed up to look plausible. This happens because AI generates answers based on inaccurate second-hand information floating around the internet rather than accessing authoritative official statistical databases directly to grasp the context. If such distorted figures were used as they are in national policymaking or in critical corporate decisions, the social cost would be hard to imagine. It would be no different from speeding in a high-performance vehicle fitted with a faulty navigation system.
The "AI-friendly metadata and ontology" project the National Data Agency is pushing forward is precisely the work of laying down precise maps and clear signposts so that AI does not lose its way. It means assigning each data set the exact context of its compilation standards, definitions and creation history, and systematically weaving a semantic network of connections among fragmented data. Once this intelligent data system is complete, AI will find official statistical databases on its own and answer with accurate context instead of wandering through inaccurate internet information. Asked about the causes of the low birth rate, for example, it would support multifaceted searches of statistical data so that the analysis goes beyond simple counts of births to connect related social indicators such as marriage rates, the age at first marriage and the burden of housing costs. Public data long confined behind ministry walls would finally begin to run on a broad highway, generating synergy.
Of course, paving this digital autobahn cannot be finished in a short time with slogans alone. The National Data Agency plans to invest about 10.8 billion won ($7.8 million) from its 2027 budget to take on the government's role as chief data officer in earnest, through work including the design of data governance, the building of AI-friendly data and the strengthening of the national data hub function. Bold and sustained financial backing is essential to expand national metadata standards across the entire government — administration, labor, diplomacy and more — starting with statistical metadata. This is a strategic decision to lay down, in advance, sturdy national data social overhead capital that Korea will benefit from for more than 30 years, rather than clinging to short-term results. On the international stage, including the Organisation for Economic Co-operation and Development's Committee on Statistics and Statistical Policy, our metadata structuring strategy has already drawn deep praise and attention as a potential global standard model.
True national competitiveness in the AI era does not stop at who has the larger model. What decides the outcome is whether a country has solid infrastructure that can reliably and accurately find and connect the data it needs at any time. Paving data highways may not be flashy, but it is the surest cornerstone Korea must lay now to become one of the world's top three AI powers.







