Ministry of Science and ICT logo./Courtesy of Ministry of Science and ICT

The Ministry of Science and ICT and the National Information Society Agency (NIA) said on the 27th that they will release 29 types of artificial intelligence (AI) training data built by five elite teams that took part in the independent AI foundation model project through AI Hub.

The released data totals about 35.44 million items and 1.56 trillion tokens. It was prepared with a 15 billion won budget for data construction and processing invested in 2025, and is enough to train large AI models with 70 billion to 80 billion parameters. In addition to large-scale pretraining data, it includes multimodal materials such as video and audio and red-teaming data to check model safety.

By team, Naver Cloud built 2.86 million public and broadcast videos, 15.5 million text items created based on them, and 12 million voice Q&A items. They can be used to develop video understanding and image generation models. Upstage released pretraining materials totaling 1 trillion tokens and 500,000 post-training data items to enhance reasoning, judgment, and execution capabilities.

SK Telecom released about 10,000 items of high-difficulty reasoning materials in mathematics, science, and law; audio and image data; and Korea-specific red-teaming data reflecting domestic laws and social values. NCAI built seven types of industry field data based on manufacturing technical documents and voice from civil complaint counseling, including long-context understanding, multi-step reasoning, and multi-turn dialogue.

LG AI Research unveiled physical AI materials filmed of more than 50 types of household chores in 50 homes in Korea. It multilayered labels more than 170,000 video clips with object, pose, and context information to enable use in research on robots and vision-language models (VLMs).

NIA and the Telecommunications Technology Association (TTA) set the scope of disclosure after verifying personal information and harmfulness, quality, and usage rights issues. Naver Cloud, Upstage, SK Telecom, and NCAI will open all materials that passed verification, and LG AI Research will statistically select and open at least 50%, the mandatory disclosure ratio.

Domestic corporations, researchers, and students can download them for free from the "independent AI model data" menu on AI Hub. Some highly sensitive materials must be used in a secure environment called the Safe Zone after a separate application. The Ministry of Science and ICT plans to disclose additional data secured during the second-stage evaluation process as soon as verification is complete.

※ This article has been translated by AI. Share your feedback here.