The public evaluation of the independent artificial intelligence (AI) foundation model project wrapped up on the 11th, moving the second evaluation process into its final stage. The Ministry of Science and ICT plans to announce the final results this month, reflecting the scores of a "public evaluation panel" made up of 200 ordinary citizens. Naver Cloud and NC AI were eliminated in the first evaluation, and in the second evaluation, one of the four teams—LG AI Research, Upstage, SK Telecom, and Motif Technologies—will be dropped.

On the 12th, ChosunBiz analyzed four Dokpamo models that the four teams released on the open-source AI platform Hugging Face through technical reports and each company's presentations. LG AI Research's "K-EXAONE 2.0," SKT's "adot X K2 (A.X K2)," Upstage's "Solar Open 2," and Motif's "Motif-3" are in the spotlight.

◇ LG and SKT bulked up; Upstage and Motif chose efficiency

LG AI Research, which ranked first in the first evaluation, more than tripled the parameters ahead of the second evaluation, from 236 billion to 750 billion, the largest in Korea. SKT also unveiled a 688 billion-parameter model, up from 519 billion in the first evaluation model. These are ultralarge models. Motif's parameters total 314 billion, and Upstage's total 250 billion.

Parameters are a key factor that determines the size of AI. They serve as the equivalent of brain cells. The more parameters, the more the model can remember and process. U.S. big tech corporations such as OpenAI, Google, and Anthropic do not disclose parameter counts, but top-tier AI models are generally built with parameters on the trillion scale and then made lighter. Moonshot AI's "Kimi K3" has 2.8 trillion parameters, and DeepSeek's "V4 Pro" has 1.6 trillion. LG AI Research said it is meaningful that "Korean researchers independently completed the entire process, from designing a 75-billion-class model to data training, distributed training, and inference environment setup," adding that it has "secured the capability to compete at the same weight class as global frontier models." SKT plans to massively scale up the successor to A.X K2 to a trillion-parameter-class model.

Upstage and Motif, by contrast, are pursuing efficiency-focused strategies that curb computation per token rather than increasing total parameters. All four models adopt a mixture-of-experts (MoE) structure that activates only some parameters suited to a user's question, and the number of such active parameters is lower for Upstage and Motif. LG AI Research has 37 billion, SKT 33 billion, Upstage 15 billion, and Motif 13.2 billion. Upstage and Motif emphasize a high-efficiency architecture that delivers large-model-level performance with fewer parameters. This reduces the number of graphics processing units (GPUs) needed to run the model, lowering introduction expense and operating costs. For example, the LG AI Research model runs on 16 Nvidia H200 GPUs (two nodes), while the Upstage model, with quantization applied, runs on just two H200 GPUs.

Bae Kyung-hoon, Deputy Prime Minister and Minister of the Ministry of Science and ICT, and attendees pose for a commemorative photo at the announcement of the Independent AI Foundation Model Project at COEX Auditorium in Gangnam-gu, Seoul, on December 30 last year. /Courtesy of News1

◇ Where the four models put their strength

It is difficult to rank the four models based on the benchmarks (performance evaluations) disclosed by each company. That is because figures vary—some measured their own and competitors' numbers in in-house environments, while others cited official releases or figures from the AI platform OpenRouter.

Picking out features from the benchmarks released by each company, LG AI Research's model has strengths in understanding long-form context, reflecting Korea's particular characteristics, and adhering to global ethical standards. It scored 94.4 on the global long-context understanding evaluation (OpenAI-MRCR) and 89.6 on the Korean long-context understanding evaluation (Ko-LongBench), showing superior performance compared to China's Zhipu GLM-5.1 (71.5 and 83.6). LG AI Research said it reflected Korea's particular characteristics in the model by operating an advisory group of 46 teachers certified in global citizenship education in cooperation with the Asia-Pacific Centre of Education for International Understanding (APCEIU) under UNESCO. On the other hand, more than tripling the parameters did not lead to performance gains across all areas. Compared with the first-evaluation model "K-EXAONE," its scores fell in areas such as the general knowledge test (MMLU-Pro, 83.8→83.5) and the math test (HMMT FEB 2026, 80.7→78.4).

The SKT model is strong in math and Korean. It scored 29 out of 42 on six problems from the "International Mathematical Olympiad (IMO) 2026," a performance equivalent to a gold medal. This year's gold medal cutoff at the IMO is 29 points. SKT said, "The math reasoning performance of an AI model is not just about solving problems; it is a proxy indicator and a core competitive edge for assessing the model's capabilities in other tasks." It also achieved top scores among comparison models such as Qwen 3.5 in the Korean knowledge evaluation (KMMLU-Pro) and Korean culture understanding evaluation (CLIcK). However, it posted the lowest score among comparison models in BrowseComp, which evaluates whether AI can search the internet on its own to find answers, indicating it lagged in locating information intricately woven across the web.

The Upstage model scored 86.8 on the Korean office task evaluation (Ko-GDPval), comparable to DeepSeek V4 Pro (86.9), which has more than six times the parameters. This benchmark has AI models directly create documents such as reports, plans, and presentation materials. The company touts performance relative to size and practical task handling as its strengths. However, Ko-GDPval is a benchmark Upstage devised itself, and it trailed DeepSeek V4 Flash in the Korean knowledge evaluation (KMMLU-Pro). The SKT model outperformed DeepSeek V4 Flash on the same benchmark.

The Motif model is strong in AI agent capabilities. It achieved top scores among comparison models such as Qwen 3.7 Max in the banking task evaluation (τ³-Banking) and the information technology (IT) incident root-cause diagnosis evaluation (ITBench). However, the Motif model includes no Korean benchmark results. Although it describes its training data as "emphasizing Korean," there is no Korean benchmark evaluation to support that.

According to the Ministry of Science and ICT, all four teams' models were selected as "Notable Models" by Epoch AI, a U.S. nonprofit AI research organization. Only three countries—the United States, China, and Korea—currently have models chosen as "Notable Models" by Epoch AI. The indicator is used as a key metric in the "AI Index" report published by the Stanford University Institute for Human-Centered Artificial Intelligence (HAI).

A professor who requested anonymity said, "When it comes to benchmark scores, models can be trained in a cramming-tutor fashion to produce good benchmark results relative to their actual capabilities," adding, "It is necessary for evaluation agencies to assess them fairly."

※ This article has been translated by AI. Share your feedback here.