Here’s a revised article in the style of CNBC, incorporating deeper analysis and professional language:
Alibaba’s AI Ambitions Escalate with Qwen3.8-Max, While DeepSeek Challenges with Aggressive Pricing
The artificial intelligence landscape is witnessing a dynamic shift as major players like Alibaba and emerging forces such as DeepSeek roll out new models, each vying for market share through distinct strategies. Alibaba has unveiled its most extensive AI model to date, Qwen3.8-Max, a behemoth boasting 2.4 trillion parameters. This launch directly confronts a market increasingly focused not just on raw power, but also on the economic viability and practical application of these advanced systems.
At the heart of Qwen3.8-Max’s architecture lies a sophisticated mixture-of-experts (MoE) design. This approach is crucial for managing the computational demands of such a colossal model. By intelligently activating only a subset of its parameters for any given request, Alibaba claims to achieve significant cost reductions and lower response latency compared to activating the entire 2.4 trillion parameters. The company states that approximately 95 billion parameters are active concurrently, a level of efficiency that is paramount for scaling AI services to a global user base.
This MoE strategy is not unique to Alibaba. DeepSeek, for instance, employs a similar, albeit smaller-scale, sparse architecture. Its latest V4-Flash model, analyzed by Artificial Analysis, features 284 billion total parameters with a more focused 13 billion active during inference. This contrasts with Moonshot AI’s Kimi K3, which offers a substantial 2.8 trillion total parameters but activates around 104 billion. The differing active parameter counts highlight a critical divergence in how these companies are approaching model efficiency and cost-effectiveness.
Qwen3.8-Max is designed for multimodal capabilities, capable of processing text, images, and video. Its impressive one-million-token context window suggests a capacity for understanding and generating complex, long-form content. Alibaba has underscored this by noting the model’s ability to complete a comprehensive software engineering project over a 16-day period, showcasing its potential for advanced task automation.
The sheer scale of Qwen3.8-Max positions it as a direct competitor to Moonshot AI’s Kimi K3. However, the competition is extending beyond model size to aggressive pricing strategies. Alibaba has set its pricing for Qwen3.8-Max at $2 per million input tokens and $6 per million output tokens. This is a direct challenge to Kimi K3’s respective pricing of $3 and $15, indicating a strategic move to capture market share through affordability.
It’s important to note that model size alone is not the sole determinant of inference cost. The underlying architecture, the ratio of active to total parameters, token consumption patterns, and the number of iterative calls required to complete a complex task all significantly influence operational expenses. This nuanced understanding is critical for businesses evaluating AI solutions.
Following its release, Qwen3.8-Max has ascended to the top position among Chinese text-based AI models on the crowdsourced comparison platform Arena.AI. While it trails some of Anthropic’s leading models in overall rankings, its strong performance in image and visual material analysis, where it secured second place on Arena.AI’s leaderboard behind an Anthropic Claude Fable 5 variant, underscores its broad utility.
DeepSeek, meanwhile, is forging a different path with its V4-Flash model, focusing aggressively on reducing inference pricing. By eschewing the ultra-large parameter counts seen in some competitors, DeepSeek is positioning V4-Flash as a highly cost-effective alternative. Artificial Analysis reports V4-Flash’s input token pricing at $0.14 per million and output at $0.28 per million. This is a significant markdown compared to other widely adopted AI systems.
Further enhancing its cost advantage, DeepSeek offers an exceptionally low cache-hit pricing of $0.003 per million tokens for the “Max Effort” version of V4-Flash. This represents a staggering 98% reduction from its standard input rate, a strategic offering for applications that benefit from repeated context processing. Cached input, which reuses previously processed information, can drastically cut down on computational overhead.
This aggressive pricing strategy is reflected in benchmark testing. Artificial Analysis estimated V4-Flash’s average cost per test at a mere three cents. This stands in stark contrast to Kimi K3’s 86 cents, OpenAI’s GPT-5.6 Sol at $1.86, and Anthropic’s Claude Fable 5 at $3.15. While a lower per-token rate is compelling, it’s crucial to remember that it doesn’t always translate to a lower overall task cost if a model requires more interactions or generates more output to achieve the desired result.
In terms of performance, Artificial Analysis assigned DeepSeek V4-Flash’s “Max Effort” reasoning version a score of 40 on its Intelligence Index, with an impressive output rate of approximately 118 tokens per second during testing.
The economic realities of deploying AI are further illustrated by Moonshot AI’s Kimi K3. While its advertised API prices are $3 per million input tokens and $15 per million output tokens, with cached input at $0.30 per million, the actual cost of completing complex tasks can diverge. On Artificial Analysis’ AA-Briefcase benchmark, designed for agentic knowledge work, Kimi K3 averaged $10.57 per task. This was driven by its generation of around 120,000 output tokens and an average of 83 turns per task.
These figures highlight that factors such as output volume and the number of model interactions can significantly inflate the total cost, moving it beyond the initial headline API rates. Kimi K3 did achieve a high score on the AA-Briefcase evaluation, ranking second behind Claude Fable 5, and scored 57 on the broader Intelligence Index, indicating its strong capabilities.
The comparative analysis between DeepSeek and Kimi K3 underscores why cost-per-task measurements are vital. Models with diverse architectures and varying usage patterns can incur substantially different compute and token costs for the same task. This level of detailed analysis is essential for businesses to make informed procurement decisions.
Adding another dimension to the competitive landscape is the growing trend of open-weight model releases. Alibaba, DeepSeek, and Moonshot AI are all providing open-weight versions of their models alongside hosted API services. This offers developers greater flexibility in deployment, allowing them to run models on their own infrastructure or through third-party cloud providers, rather than being solely reliant on a developer’s inference service.
DeepSeek V4-Flash is available as an open-weight model under an MIT license through Hugging Face, while Kimi K3 also offers an open-weight option under Moonshot AI’s proprietary license. While deployment costs will still vary based on internal infrastructure, this approach decouples access to the model from a single hosted API. This contrasts with the proprietary, closed-weight models favored by industry giants like OpenAI, Anthropic, and Google.
Lian Jye Su, chief analyst at Omdia, notes that for many enterprise workloads, the decision-making process extends beyond simply identifying the highest-performing system. “Many business workflows do not need the industry’s very best model,” Su commented. “They need models that are good enough, affordable, transparent, and accessible, and open-weight models help meet that demand.” This sentiment indicates a maturing market where practical considerations like cost and accessibility are increasingly influencing AI adoption.
Original article, Author: Samuel Thompson. If you wish to reprint this article, please indicate the source:https://aicnbc.com/24441.html