Korea Opens 35 Million AI Training Data Records for Free Use

Five Teams in Sovereign AI Project Release 29 Datasets About 35.44 Million Records, Equal to 1.56 Trillion Tokens Covers Video, Speech, AI Agents and Robotics NAVER, Upstage, SK Telecom and NC AI Release All Verified Data LG AI Research Discloses Selected Data Above the 50% Minimum

Technology|
|
By Seo Ji-hyewise@sedaily.com
||

The South Korean government will open about 35.44 million AI training data records secured through its sovereign artificial intelligence foundation model project to the private sector free of charge. The aim is to lower barriers to AI development by letting domestic companies and researchers tap the data-building know-how of firms that took part in developing the country's flagship AI models.

The Ministry of Science and ICT and the National Information Society Agency said on the 27th that they will release 29 datasets on AI Hub, built by the five elite teams that participated in the first-stage evaluation of the sovereign AI foundation model project — NAVER Cloud, Upstage, SK Telecom (017670.KS), NC AI and LG AI Research. NAVER Cloud and NC AI were later dropped in subsequent stage evaluations, but their data remained subject to mandatory disclosure because it was secured with government funding.

The data released totals about 35.44 million records and 11.3 terabytes, equal to roughly 1.56 trillion tokens in training terms. In theory, that is enough to train a large AI model with 70 billion to 80 billion parameters. NAVER Cloud, Upstage, SK Telecom, NC AI and LG AI Research each prepared the data to fit their own model development strategies, drawing on the 15 billion won the government allocated in 2025 for data construction and processing.

Image created with ChatGPT - Seoul Economic Daily Technology News from South Korea
Image created with ChatGPT

NAVER Cloud built datasets of publicly available and broadcast video, scene-by-scene captions, and speech-based question-and-answer data. These can be used to develop video understanding and image generation models. Upstage is releasing 1 trillion tokens of pretraining data along with post-training data for AI agents that strengthens reasoning, judgment and execution capabilities.

SK Telecom is providing advanced reasoning material in mathematics, science and law, as well as red-teaming data reflecting universal values in Korean society. NC AI is offering long-context understanding, multi-turn dialogue and multimodal datasets built from manufacturing technical documents and public complaint consultation records. LG AI Research is releasing training data for humanoid robots, filmed in actual Korean households covering more than 50 types of household chores.

The disclosures stem from the participation terms of the sovereign AI foundation model program. Elite teams must release at least 50% of the material secured with the government's data construction and processing budget. NAVER Cloud, Upstage, SK Telecom and NC AI decided to open all data that cleared quality verification. LG AI Research is releasing material extracted statistically at fixed intervals across each dataset to meet the mandatory 50% threshold. The selection process was verified by the Telecommunications Technology Association.

The government plans to spread the AI development assets accumulated through the project across the domestic ecosystem rather than leaving them with individual companies. The intent is to allow startups, universities and researchers that cannot secure and refine high-quality training data on their own to use material employed in developing the country's flagship AI in building models and improving performance. As a result, the scope of disclosure has widened from text-centered pretraining material to multimodal data covering video, speech and images, as well as AI agents, safety and physical AI.

Any domestic company, researcher or student can search, download and use the released data free of charge under the "Sovereign AI Model Data" category on AI Hub. Some material requiring personal data protection or security management, however, must be requested separately on the website and used through a secured online development environment called the Safety Zone. The ministry said data newly built during the second-stage evaluation will also be released after quality verification.

Original reporting by Seo Ji-hye for Seoul Economic Daily.

AI-translated from Korean. Quotes from foreign sources are based on Korean-language reports and may not reflect exact original wording.

Watch · Seoul Economic Daily

More →
3:02

AI KEY

Preview
Korean Corporate Intelligence HubKOSPI · KOSDAQ · 12 sectors

A live, cap-weighted view of every KOSPI and KOSDAQ sector, with same-day Korean reporting distilled by company — built for foreign investors, correspondents and analysts who need to scan Korea before the next session.

Korea Company Atlas

Preview
Market Ontology · The Feedback LoopKFTC 2025 · 92 groups · 121,954 articles

An English ontology of the Korean market — how companies, the media, the government and the National Assembly move each other in a loop. Korea's named controlling persons and designated business groups are a mechanism, not a risk to be priced blind.

SIGNAL

Now live
English Edition · Capital MarketsM&A · IPO · PE · Fund Flows

SIGNAL English Edition is live — Korea's deal desk reporting in English. M&A, IPOs, private equity and fund flows, covered daily for global institutional investors. Browse free; subscriber-only scoops at the 50% intro rate.