About this model
DeepSeek V4 Flash 0731 is the efficiency-oriented member of DeepSeek's V4 generation: a sparse Mixture-of-Experts model with 284B total parameters and roughly 13B activated per token, paired with a one-million-token context window. It sits below the larger DeepSeek V4 Pro, which shares the same million-token context but is aimed at the heaviest knowledge and agentic workloads.
Architecturally, the V4 line departs from DeepSeek V3.2 through a hybrid attention design that combines compressed sparse attention with heavily compressed attention. DeepSeek's own model card describes this as substantially reducing per-token inference FLOPs and KV cache footprint at million-token context relative to V3.2 — the efficiency recipe that lets the Flash tier serve very long inputs at high throughput. The card also outlines a two-stage post-training pipeline: domain-specific expert models trained with supervised fine-tuning and reinforcement learning, then consolidated into a single model via on-policy distillation.
Compared with the earlier April 2026 DeepSeek V4 Flash release, the 0731 build is a dated refresh within the same family: parameter counts and context window are unchanged, so differences come from updated training rather than a new architecture. Independent evaluation from Artificial Analysis places the reasoning, maximum-effort configuration at 50 on its Intelligence Index.
Catalog capabilities include reasoning, code optimization, function calling and web search, suiting long-context agents, repository-scale coding and high-volume batch pipelines. An encrypted-serving variant is listed separately as DeepSeek V4 Flash.
This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.
Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 6h ago