DeepSeek-V4.1-Flash Launches on Qwen AI Platform, API and Token Plan Opened Simultaneously - AIBase
Article excerpt
Highlighted: the sentence this signal was extracted from
DeepSeek-V4.1-Flash has recently been launched on the Qwen AI platform, with API services and Token Plan now available. Developers can integrate the model into their systems via standard APIs, or directly use the Token Plan in tools such as Qoder, Qwen APP, and Codex for code, documentation, visual understanding, and agent tasks. According to the platform announcement, this model is DeepSeek's new lightweight flagship, featuring a MoE architecture with 552B total parameters, using a Causal-Encoder-Decoder asymmetric structure, with input activation of about 8B and output activation of about 16B; it natively supports text and image understanding, with a maximum context length of 1 million tokens, and a maximum output of approximately 393K. The official stated that it has improved significantly over its predecessor in several Agent and code benchmarks, achieving high throughput and low latency with lower activation parameters. The cost side is a major focus of this release. The new generation of cache compression reduces the KV Cache demand for HBM to one-quarter of the previous generation and for SSD storage to one-eighth, saving more resources for long contexts and multi-turn tool calls. The Qwen page provides time-based pricing: 1 yuan per million tokens during off-peak hours for input and 4 yuan per million tokens for output; 2 yuan and 8 yuan respectively during peak...
Keep reading with a free account
The rest of this article, and every signal for DeepSeek, is in your free account.