Noise continues around the independent AI foundation model project the government is pushing as it says it will pick a "national AI." Motif Technologies was eliminated despite earning the top score on the global AAII benchmark in the second round. Questions are being raised about whether it is reasonable for a team that joined late through a "losers' revival" and ranked No. 1 in performance to be cut, and there is criticism that the Ministry of Science and ICT's belated disclosure of detailed scores was not handled smoothly.
◇ A review that knocks out the global benchmark No. 1
According to the Ministry of Science and ICT on the 28th, Motif ranked No. 1 in the AAII benchmark review (out of 25 points) but placed fourth in the NIA benchmark review (out of 15 points), expert review (out of 35 points), and user review (out of 25 points), finishing fourth overall and being eliminated. SK Telecom ranked first overall with 70.6 points, Upstage second with 69.9, and LG AI Research Institute third with 69.0. Motif received 65.8.
The AAII benchmark is a new evaluation method the Ministry of Science and ICT introduced to ensure reliability and fairness in the second-round review. At the time, the ministry said it was "a highly credible benchmark on which even leading global big techs find it difficult to score high." Motif scored 47 on AAII, ranking first among the four participating models and tenth globally.
While there were questions about whether the benchmark can truly assess a model's performance and controversy over whether Motif engaged in so-called "benchmaxing" to boost its benchmark score, Motif's claim that there is a problem with how the benchmark weight was calculated is also persuasive. Anthropic's Opus 5, a top-tier global model, has an AAII score of 63, which converts to only 15.75 on a 25-point scale, meaning actual performance gaps between models are not sufficiently reflected in the evaluation score. In the second-round review, AAII benchmark scores were Motif 11.9, Upstage 9.4, SK Telecom 8.8, and LG AI Research Institute 7.8. NIA benchmark scores, out of 15, ranged from 12.7 to 13.4, while AAII benchmark scores, out of 25, ranged from 7.8 to 11.9. The AAII reviewing body also said there was no benchmaxing by participants.
LG AI Research Institute, which swept all three first-round categories—benchmark, expert, and user—and took a commanding lead with 90.2 points (the five-team average was 79.7), fell to third in the second round, showing how rankings can be completely reshuffled by the government's change in evaluation method. That is because LG AI Research Institute received the lowest score on the newly added AAII benchmark in the second round, pushing it down the rankings.
With the final scores of the top three bunched tightly between 69.0 and 70.6, the scoring tendencies of individual reviewers had a sizable impact on the rankings. In the detailed expert review scores (out of 35) released by the Ministry of Science and ICT the previous day, "Commissioner7" among the 10 reviewers gave Upstage 30 points, LG AI Research Institute 24, SK Telecom 20, and Motif 17. This reviewer showed a tendency to give a wide spread of scores, with a 13-point gap between the lowest and highest. By contrast, the other reviewers had an average gap of only 4.6 points between their lowest and highest scores. The government adopted a method that averages the remaining scores after excluding the highest and lowest, and while "Commissioner7"'s scores for LG AI Research Institute, SK Telecom, and Motif were the lowest among the 10 and thus excluded, the 30-point score for Upstage was not the highest and was therefore counted.
Recalculating the expert review using the same method but excluding "Commissioner7," whose variance was unusually large, narrows the overall score gap between second-place Upstage and third-place LG AI Research Institute from 0.9 points to the 0.1-point range.
◇ The process wasn't smooth either
The process also wasn't smooth. When the Ministry of Science and ICT announced the second-round results on Aug. 18, it disclosed only which teams advanced or were eliminated, without releasing the advancing teams' rankings or detailed scores. After criticism that this was a taxpayer-funded program and even the rankings were being withheld, on the 20th the ministry partially disclosed only category leaders, citing a "stigma effect" on lower-ranked teams. Full scores across all categories were revealed only after Motif issued a statement and demanded full disclosure. In the first round, Naver Cloud was eliminated for failing to meet the "independence" criterion, and the subsequently unscheduled losers' revival was held, sparking controversy over the policy's credibility and fairness. Questions of fairness were amplified by conflict-of-interest allegations involving Upstage and Ha Jung-woo, former Blue House presidential secretary for AI future planning.
There is also a fundamental critique that the independent AI foundation model project's goals and evaluation metrics are misaligned. The project began with a target of achieving at least 95% of frontier AI model performance, but in the second round the weighting favored human-scored expert and user reviews (60 points) over benchmarks (40 points). Industry figures ask whether boosting usability and spreading services should be part of the "Everyone's AI" program, not the core yardstick for a project that is supposed to compete on technical prowess.
The government acknowledges the limits of the approach. Bae Kyung-hoon, Deputy Prime Minister and Minister of the Ministry of Science and ICT, said, "It is time to fundamentally consider whether evaluation and competition methods established a year and a half ago remain valid amid AI trends that change every three to six months," adding, "To build a world-class model that anyone can recognize, we will review supplementing the way we concentrate government investment and our evaluation system."
Kim Jang-hyun, a professor at Sungkyunkwan University, said, "No matter how advanced a model is, if user acceptance is low, the significance of its performance is diminished," but added, "There is an enormous global race to raise large language model (LLM) performance even slightly. It may look like a simple numbers game, but in fact the scores compare tests across various dimensions, so benchmarks aren't that simple." He continued, "The government should further refine its screening criteria for the usability component that sparked controversy this time."