The Information reported on the 13th that a study found OpenAI and Anthropic's artificial intelligence (AI) models are actually cheaper to use for some tasks than China's open-weight (open-weights) AI models.
AlphaSense, an AI-based search and market research firm, released study results earlier this month showing that the closed models from OpenAI and Anthropic produced better answers for some tasks at lower expense than China's open-weight models.
AlphaSense tested the answer quality and expense of nine models using 245 finance analysis questions based on earnings releases, investor conference call transcripts, and securities filings.
As a result, OpenAI's "GPT-5.6 Sol" and Anthropic's "Opus 4.8" delivered higher-quality answers at lower expense than China Moonshot's "Kimi K3" and Zhipu (now Z.ai)'s "GLM-5.2."
In general, looking only at token prices, the China models are cheaper. Per 1 million output tokens, Kimi K3 is $15, Opus 4.8 is $25, and GPT-5.6 Sol is $30.
However, in the case of Kimi K3, while the per-token rate is lower, it consumes more tokens to produce an answer. As a result, AlphaSense said the actual total expense per question was higher than GPT-5.6 Sol.
GPT-5.6 Sol's median total expense was 13% lower than Kimi K3's, while its quality score was 20% higher. Opus 4.8 outperformed Kimi K3 by 13% in quality at half the expense. Zhipu's GLM-5.2 had about twice the expense of Kimi K3 but lower quality.
The study was released as Kimi K3 showed performance on par with U.S. closed AI models on some benchmarks last month, prompting speculation that OpenAI and Anthropic's price competitiveness could weaken.
AlphaSense CEO Jack Kokko said, "Models that look expensive when judged only by price per token actually turned out to use tokens more efficiently, incurring less expense."
Earlier, at the end of last month, OpenAI board chair Bret Taylor said in an interview with CNBC that top-tier models are cheaper and more efficient than open-source models because they use fewer tokens when answering questions or doing tasks.
However, results can vary by evaluation method. Artificial Analysis, an AI model analytics firm, assessed that when compared on the same basis, the expense of Kimi K3 and GLM-5.2 was much cheaper than U.S. models. In other words, conclusions about models' price competitiveness can change depending on which tasks are targeted and how expense is calculated.
Kokko said the ideal approach is to mix multiple models. AlphaSense applies this principle in its own search service. Depending on the nature of the question, it lowers expense with an internal harness (engine) that assigns a high-performing model for planning and a cheaper model for execution.