{"id":23000,"date":"2026-08-04T00:54:11","date_gmt":"2026-08-04T00:54:11","guid":{"rendered":"https:\/\/scannn.com\/deepseek-ai-deepseek-v4-flash-0731-%c2%b7-hugging-face\/"},"modified":"2026-08-04T00:54:11","modified_gmt":"2026-08-04T00:54:11","slug":"deepseek-ai-deepseek-v4-flash-0731-%c2%b7-hugging-face","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/deepseek-ai-deepseek-v4-flash-0731-%c2%b7-hugging-face\/","title":{"rendered":"deepseek-ai\/DeepSeek-V4-Flash-0731 \u00b7 Hugging Face"},"content":{"rendered":"\n<div><!--[-1--><!--]--> <!----><\/p>\n<div align=\"center\">\n  \n<\/div>\n<hr\/>\n<div align=\"center\" style=\"line-height: 1;\">\n  <a href=\"https:\/\/www.deepseek.com\/\" style=\"margin: 2px;\" rel=\"nofollow\"><br \/>\n    <img decoding=\"async\" alt=\"Homepage\" src=\"https:\/\/github.com\/deepseek-ai\/DeepSeek-V2\/blob\/main\/figures\/badge.svg?raw=true\" style=\"display: inline-block; vertical-align: middle;\"\/><br \/>\n  <\/a><br \/>\n  <a href=\"https:\/\/chat.deepseek.com\/\" style=\"margin: 2px;\" rel=\"nofollow\"><br \/>\n    <img decoding=\"async\" alt=\"Chat\" src=\"https:\/\/img.shields.io\/badge\/&#x1f916;%20Chat-DeepSeek%20V4-536af5?color=536af5&amp;logoColor=white\" style=\"display: inline-block; vertical-align: middle;\"\/><br \/>\n  <\/a>\n<\/div>\n<div align=\"center\" style=\"line-height: 1;\">\n  <a href=\"https:\/\/huggingface.co\/deepseek-ai\" style=\"margin: 2px;\"><br \/>\n    <img decoding=\"async\" alt=\"Hugging Face\" src=\"https:\/\/img.shields.io\/badge\/%F0%9F%A4%97%20Hugging%20Face-DeepSeek%20AI-ffc107?color=ffc107&amp;logoColor=white\" style=\"display: inline-block; vertical-align: middle;\"\/><br \/>\n  <\/a><br \/>\n  <a href=\"https:\/\/twitter.com\/deepseek_ai\" style=\"margin: 2px;\" rel=\"nofollow\"><br \/>\n    <img decoding=\"async\" alt=\"Twitter Follow\" src=\"https:\/\/img.shields.io\/badge\/Twitter-deepseek_ai-white?logo=x&amp;logoColor=white\" style=\"display: inline-block; vertical-align: middle;\"\/><br \/>\n  <\/a>\n<\/div>\n<div align=\"center\" style=\"line-height: 1;\">\n  <a href=\"https:\/\/huggingface.co\/deepseek-ai\/LICENSE\" style=\"margin: 2px;\" rel=\"nofollow\"><br \/>\n    <img decoding=\"async\" alt=\"License\" src=\"https:\/\/img.shields.io\/badge\/License-MIT-f5de53?&amp;color=f5de53\" style=\"display: inline-block; vertical-align: middle;\"\/><br \/>\n  <\/a>\n<\/div>\n<p align=\"center\">\n  <a href=\"https:\/\/arxiv.org\/abs\/2606.19348\" rel=\"nofollow\"><b>Technical Report<\/b><\/a>\n<\/p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"introduction\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#introduction\" rel=\"nofollow\"><\/p>\n<p>\t<\/a><br \/>\n\t<span><br \/>\n\t\tIntroduction<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p><strong>DeepSeek-V4-Flash-0731<\/strong> is the official release of <strong>DeepSeek-V4-Flash<\/strong>, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Flash-DSpark\" rel=\"nofollow\">DeepSeek-V4-Flash-DSpark<\/a>, i.e. it comes with a speculative decoding module attached.<\/p>\n<p>DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.<\/p>\n<div align=\"center\">\n<div class=\"max-w-full overflow-auto\">\n<table>\n<thead>\n<tr>\n<th align=\"left\">Benchmark<\/th>\n<th align=\"center\">DeepSeek-V4-Flash-0731<\/th>\n<th align=\"center\">DeepSeek-V4-Flash (Preview)<\/th>\n<th align=\"center\">DeepSeek-V4-Pro (Preview)<\/th>\n<th align=\"center\">GLM-5.2<\/th>\n<th align=\"center\">Opus-4.8<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td align=\"left\">Terminal Bench 2.1<\/td>\n<td align=\"center\">82.7<\/td>\n<td align=\"center\">61.8<\/td>\n<td align=\"center\">72.1<\/td>\n<td align=\"center\">81.0<\/td>\n<td align=\"center\">85.0<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">NL2Repo<\/td>\n<td align=\"center\">54.2<\/td>\n<td align=\"center\">39.4<\/td>\n<td align=\"center\">38.5<\/td>\n<td align=\"center\">48.9<\/td>\n<td align=\"center\">69.7<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">Cybergym<\/td>\n<td align=\"center\">76.7<\/td>\n<td align=\"center\">38.7<\/td>\n<td align=\"center\">52.7<\/td>\n<td align=\"center\">&#8211;<\/td>\n<td align=\"center\">83.1<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">DeepSWE<\/td>\n<td align=\"center\">54.4<\/td>\n<td align=\"center\">7.3<\/td>\n<td align=\"center\">12.8<\/td>\n<td align=\"center\">46.2<\/td>\n<td align=\"center\">58.0<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">Toolathlon-Verified<\/td>\n<td align=\"center\">70.3<\/td>\n<td align=\"center\">49.7<\/td>\n<td align=\"center\">55.9<\/td>\n<td align=\"center\">59.9<\/td>\n<td align=\"center\">76.2<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">Agents&#8217; Last Exam<\/td>\n<td align=\"center\">25.2<\/td>\n<td align=\"center\">15.8<\/td>\n<td align=\"center\">16.5<\/td>\n<td align=\"center\">23.8<\/td>\n<td align=\"center\">25.7<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">AutomationBench Public<\/td>\n<td align=\"center\">25.1<\/td>\n<td align=\"center\">10.8<\/td>\n<td align=\"center\">12.8<\/td>\n<td align=\"center\">12.9<\/td>\n<td align=\"center\">27.2<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">DSBench-FullStack \u2020<\/td>\n<td align=\"center\">68.7<\/td>\n<td align=\"center\">37.0<\/td>\n<td align=\"center\">41.8<\/td>\n<td align=\"center\">61.8<\/td>\n<td align=\"center\">71.6<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">DSBench-Hard \u2020<\/td>\n<td align=\"center\">59.6<\/td>\n<td align=\"center\">25.8<\/td>\n<td align=\"center\">31.1<\/td>\n<td align=\"center\">54.5<\/td>\n<td align=\"center\">71.7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<\/div>\n<p>Notes:<\/p>\n<ol>\n<li>For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the <code>max<\/code> reasoning effort level with <code>temperature = 1.0, top_p = 0.95<\/code>.<\/li>\n<li>\u2020 DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.<\/li>\n<\/ol>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"chat-template\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#chat-template\" rel=\"nofollow\"><\/p>\n<p>\t<\/a><br \/>\n\t<span><br \/>\n\t\tChat Template<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>This release does not include a Jinja-format chat template. Instead, we provide a dedicated <code>encoding<\/code> folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model&#8217;s text output. Please refer to the <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Flash-0731\/blob\/main\/encoding\/README.md\" rel=\"nofollow\"><code>encoding<\/code><\/a> folder for full documentation.<\/p>\n<p>The <code>reasoning_effort<\/code> parameter now supports three levels \u2014 <code>low<\/code>, <code>high<\/code>, and <code>max<\/code> \u2014 which control how much deliberation the model spends before answering.<\/p>\n<p>A brief example:<\/p>\n<pre><code class=\"language-python\"><span class=\"hljs-keyword\">from<\/span> encoding_dsv4 <span class=\"hljs-keyword\">import<\/span> encode_messages, parse_message_from_completion_text\n\nmessages = [\n    {<span class=\"hljs-string\">\"role\"<\/span>: <span class=\"hljs-string\">\"user\"<\/span>, <span class=\"hljs-string\">\"content\"<\/span>: <span class=\"hljs-string\">\"hello\"<\/span>},\n    {<span class=\"hljs-string\">\"role\"<\/span>: <span class=\"hljs-string\">\"assistant\"<\/span>, <span class=\"hljs-string\">\"content\"<\/span>: <span class=\"hljs-string\">\"Hello! I am DeepSeek.\"<\/span>, <span class=\"hljs-string\">\"reasoning_content\"<\/span>: <span class=\"hljs-string\">\"thinking...\"<\/span>},\n    {<span class=\"hljs-string\">\"role\"<\/span>: <span class=\"hljs-string\">\"user\"<\/span>, <span class=\"hljs-string\">\"content\"<\/span>: <span class=\"hljs-string\">\"1+1=?\"<\/span>}\n]\n\n\nprompt = encode_messages(messages, thinking_mode=<span class=\"hljs-string\">\"thinking\"<\/span>, reasoning_effort=<span class=\"hljs-string\">\"max\"<\/span>)\n\n\n<span class=\"hljs-keyword\">import<\/span> transformers\ntokenizer = transformers.AutoTokenizer.from_pretrained(<span class=\"hljs-string\">\"deepseek-ai\/DeepSeek-V4-Flash-0731\"<\/span>)\ntokens = tokenizer.encode(prompt)\n<\/code><\/pre>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"how-to-run-with-vllm\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#how-to-run-with-vllm\" rel=\"nofollow\"><\/p>\n<p>\t<\/a><br \/>\n\t<span><br \/>\n\t\tHow to Run with vLLM<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>DSpark speculative decoding is enabled with a single flag \u2014 add &#8211;speculative-config with method: dspark to your vLLM launch command:<\/p>\n<p><code>--speculative-config '{\"method\":\"dspark\",\"num_speculative_tokens\":7,\"draft_sample_method\":\"greedy\"}'<\/code><\/p>\n<p>For example, the command below serves the model with vLLM on a single 4\u00d7GB300 node.<br \/>\nSee the <a href=\"https:\/\/recipes.vllm.ai\/deepseek-ai\/DeepSeek-V4-Flash?hardware=b300&amp;features=tool_calling,reasoning\" rel=\"nofollow\">vLLM recipe<\/a> for detailed instructions and other hardware configurations.<\/p>\n<pre><code class=\"language-bash\">vllm serve deepseek-ai\/DeepSeek-V4-Flash-0731 \\\n  --trust-remote-code --kv-cache-dtype fp8 --block-size 256 \\\n  --data-parallel-size 4 --enable-expert-parallel \\\n  --moe-backend deep_gemm_mega_moe \\\n  --attention-config <span class=\"hljs-string\">'{\"use_fp4_indexer_cache\": true}'<\/span> \\\n  --speculative-config <span class=\"hljs-string\">'{\"method\":\"dspark\",\"num_speculative_tokens\":7,\"draft_sample_method\":\"greedy\"}'<\/span>\n<\/code><\/pre>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"how-to-run-with-sglang\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#how-to-run-with-sglang\" rel=\"nofollow\"><\/p>\n<p>\t<\/a><br \/>\n\t<span><br \/>\n\t\tHow to Run with SGLang<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Enable DSpark with <code>--speculative-algorithm DSPARK<\/code> and do not set a separate <code>--speculative-draft-model-path<\/code> as the target and draft weights therefore come from the same checkpoint.<br \/>\nSee the <a href=\"https:\/\/docs.sglang.io\/cookbook\/autoregressive\/DeepSeek\/DeepSeek-V4#hw=gb300&amp;variant=flash-official&amp;quant=fp4&amp;strategy=low-latency&amp;nodes=single\" rel=\"nofollow\">SGLang cookbook<\/a> for detailed instructions, benchmarks and other hardwares configurations.<\/p>\n<pre><code class=\"language-bash\">sglang serve \\\n  --trust-remote-code \\\n  --model-path deepseek-ai\/DeepSeek-V4-Flash-0731 \\\n  --tp 4 \\\n  --moe-runner-backend flashinfer_mxfp4 \\\n  --speculative-algorithm DSPARK \\\n  --mem-fraction-static 0.90 \\\n  --chunked-prefill-size 4096 \\\n  --swa-full-tokens-ratio 0.1 \\\n<\/code><\/pre>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"how-to-run-locally\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#how-to-run-locally\" rel=\"nofollow\"><\/p>\n<p>\t<\/a><br \/>\n\t<span><br \/>\n\t\tHow to Run Locally<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>Please refer to the <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Flash-0731\/blob\/main\/inference\/README.md\" rel=\"nofollow\">inference<\/a> folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.<\/p>\n<p>For local deployment, we recommend setting the sampling parameters to <code>temperature = 1.0<\/code>, with <code>top_p = 0.95<\/code> for agentic scenarios and <code>top_p = 1.0<\/code> otherwise. For the <code>high<\/code> and <code>max<\/code> reasoning effort levels, we recommend a maximum output length of <strong>384K<\/strong> tokens.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"license\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#license\" rel=\"nofollow\"><\/p>\n<p>\t<\/a><br \/>\n\t<span><br \/>\n\t\tLicense<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>This repository and the model weights are licensed under the <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Flash-0731\/tree\/main\/LICENSE\" rel=\"nofollow\">MIT License<\/a>.<\/p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"citation\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#citation\" rel=\"nofollow\"><\/p>\n<p>\t<\/a><br \/>\n\t<span><br \/>\n\t\tCitation<br \/>\n\t<\/span><br \/>\n<\/h2>\n<pre><code>@misc{deepseekai2026deepseekv4,\n      title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},\n      author={DeepSeek-AI},\n      year={2026},\n}\n<\/code><\/pre>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"contact\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#contact\" rel=\"nofollow\"><\/p>\n<p>\t<\/a><br \/>\n\t<span><br \/>\n\t\tContact<br \/>\n\t<\/span><br \/>\n<\/h2>\n<p>If you have any questions, please raise an issue or contact us at <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Flash-0731\/blob\/main\/service@deepseek.com\" rel=\"nofollow\">service@deepseek.com<\/a>.<\/p>\n<p><!----><\/div>\n<p><script async src=\"\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script><br \/>\n<br \/><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Flash-0731?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Technical Report Introduction DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23001,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23000","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23000","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23000"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23000\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23001"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23000"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23000"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23000"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}