DeepSeekDeepSeek·💬 Text Generation·New

DeepSeek V4 Pro 0813

ReasoningCodeFunction CallingWeb Searchprivate
🧠 Try in Intelligence →Try on Venice.ai ↗
Quick reference
DeepSeek V4 Pro 0813 — TLDR
  • 🧠 1.6T-parameter Mixture-of-Experts flagship, roughly 49B parameters activated per token
  • 📏 One-million-token context window, standard across DeepSeek's V4 services
  • 🔧 Hybrid attention: Compressed Sparse Attention plus Heavily Compressed Attention
  • ⚡ DeepSeek reports 27% of V3.2's per-token FLOPs at 1M context
  • 🎯 Aimed at reasoning, coding, and long-horizon agentic workflows
  • 💬 Thinking and non-thinking modes selectable per request
  • 🔧 Function calling and web search enabled in this catalog deployment
  • 🏢 From DeepSeek in Hangzhou; weights published on Hugging Face
💰 Pricing
$1.73 / $4.95
per 1M · input / output
📏 Context
1M tokens
📅 On Venice since
Aug 14, 2026
3 days ago
Provider

DeepSeek is a Chinese artificial intelligence company specializing in large language model development, founded in July 2023 by Liang Wenfeng. Based in Hangzhou, Zhejiang, the company is backed by High-Flyer, a prominent Chinese hedge fund also co-founded by…

Read full profile →
7 models on Venice
7 text
Since Dec 4, 2025

About this model

DeepSeek V4 Pro 0813 is the dated August 2026 build of DeepSeek's flagship V4 Pro line: a Mixture-of-Experts model with roughly 1.6 trillion total parameters, about 49 billion activated per token, and a one-million-token context window, per DeepSeek's model card. It succeeds the April 2026 preview release DeepSeek V4 Pro, which introduced the same parameter scale and context length; 0813 is the later iteration of that endpoint.

Architecturally, the V4 generation departs from DeepSeek V3.2 with a hybrid attention design combining Compressed Sparse Attention and Heavily Compressed Attention. On DeepSeek's own reported figures, at one-million-token context V4 Pro needs about 27% of the single-token inference FLOPs and 10% of the KV cache of V3.2 — the stated reason million-token context became the default across DeepSeek's services rather than a premium tier. Post-training follows a two-stage recipe: domain-specific experts trained separately with supervised fine-tuning and reinforcement learning, then consolidated into a single model through on-policy distillation.

Within the family, Pro is the heavyweight tier. The lighter siblings — DeepSeek V4 Flash 0423, DeepSeek V4 Flash 0731, its low-latency DeepSeek V4 Flash 0731 Fast variant, and the encrypted DeepSeek V4 Flash deployment — use a 284B-parameter, 13B-active configuration at the same context length. DeepSeek notes that Flash's smaller scale trails Pro on knowledge-heavy tasks and the most complex agentic workflows.

The model exposes thinking and non-thinking modes, and this deployment adds function calling and web search. It is text-only: suited to long-document analysis, software engineering, tool use, and multi-step agents rather than image understanding.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 2d ago