Motif Technologies, which was eliminated in the second-round evaluation for the independent artificial intelligence (AI) foundation model project (Dokpamo), demanded disclosure of the detailed evaluation criteria and a reconsideration.

Motif on the 27th said it formally filed an objection regarding the second-stage Dokpamo selection results and released a statement of position and the items of objection.

Motif demanded: disclosure of detailed scores and evaluation content by overall evaluation item and by company; presentation of the basis for how benchmark evaluation scores were allocated; disclosure of detailed criteria and grounds for judgment in the expert evaluation; and disclosure of the user evaluation methodology and whether a blind evaluation was applied.

Motif said that while it respected and accepted the evaluation results themselves, media reports said the Ministry of Science and ICT explained that Motif 3 scored high only on the AAII benchmark but had issues in the expert and user evaluations, and that in remarks by a senior ministry official reported by some outlets, there was even an assessment that Motif 3's actual reasoning performance was insufficient compared with its benchmark score and that its usability and practicality were lacking.

It continued, We believe this goes beyond the question of which company was selected and which was eliminated, and raises fundamental questions about the essence of the Dokpamo project and the direction it should take going forward, and we also came to think there may be some misunderstandings about the technical direction we have pursued and its outcomes, so we respectfully request the detailed disclosure of the evaluation scores and a reconsideration.

Motif's Motif 3 scored 47 on AAII, a global comprehensive AI performance index. This is higher than the models from Upstage (37), SK Telecom (35), and LG AI Research Institute (31).

Motif said, In LG's case, which had the lowest AAII score, it more than doubled the model size compared with the first round, but improvement in its benchmark score was minimal, and added, It is difficult to understand the specific grounds for the result that the lowest-performing model ranked first in the expert evaluation. It also said, During the expert evaluation process, we received in writing a question such as, 'Within the consortium's lead and participating entities, is there a company in which foreign capital or an overseas corporation participates as a major shareholder with voting rights?,' the purpose of which is difficult to understand.

Motif raised issues with the score allocation and the expert and user evaluation methods. The Dokpamo second-round evaluation consisted of 40 points for benchmarks (25 for AAII and 15 for NIA), 35 for expert evaluation, and 25 for user evaluation, and in the AAII benchmark Motif ranked first with 11.9 out of 25.

Motif said that while global frontier models score 63 on AAII, converting this to a 25-point scale yields only 15.75, making it difficult for actual performance gaps between models to be sufficiently reflected in the evaluation score.

Regarding the user evaluation as well, it said, It is hard to rule out the possibility that brand awareness or preconceptions affected the evaluation, and requested disclosure of whether a blind evaluation method was applied and the detailed results.

Motif, however, said, The objection is not to argue that we must be selected for Dokpamo, and added, Regardless of the outcome of the objection, we will not participate in the third round of Dokpamo.

※ This article has been translated by AI. Share your feedback here.