DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis (artificialanalysis.ai)
476 points by theanonymousone 12 hours ago
pmxi 5 hours ago
I have updated OpenAI's chart[1] from yesterday to include one more datapoint: DeepSeek V4 Flash 0731. It's on the frontier.
https://files.parasmittal.com/openai_aa_luna_dsflash.svg
1: https://openai.com/index/advancing-the-price-performance-fro...
bwfan123 5 hours ago
> For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework
So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the same model with fireworks/openrouter, with zdr thrown in, token costs ratchet up with no explanation. Likely that the model is subsidized for gathering usage data. I am waiting for the day I can run this locally.
L8D 5 hours ago
If you're not paying for the model through Open router, then which provider are you buying it from? I thought that if that provider was listed on open router it would be the same price as going directly to the provider? Are you buying from a provider that isn't listed on open router already?
bwfan123 5 hours ago
> Are you buying from a provider that isn't listed on open router already
I am using openrouter with zdr guardrail which routes to any provider that supposedly doesnt train on user data. I also use fireworks (directly, not via openrouter) which is a provider promising zdr and has a bunch of open weights models. My issue is that these zdr providers dont transparently disclose caching/tokens etc and so they end up being far more expensive than directly using DS.
x312 2 hours ago
segmondy 5 hours ago
They announced it already, read the tech report.
"For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95."
bwfan123 5 hours ago
> They announced it already
I meant something I could download and run.
ignoramous 3 hours ago
scosman 6 hours ago
So GLM 5.2/Gemini 3.6 level intelligence for $0.28/m output. And their updated Pro model coming soon....
Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB.
luckydata 5 hours ago
I would like to see your "home"
cmrdporcupine 5 hours ago
Two (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest.
But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
ycui7 4 hours ago
sourcecodeplz an hour ago
i am scared for the PRO model
maybe it is Fable level
segmondy 6 hours ago
Q8 is ~ 151gb
scosman 6 hours ago
thanks, corrected! I looked at Q3
0cf8612b2e1e 6 hours ago
Somewhat relatedly, how do the economics for Huggingface work? They must be hosting petabytes of models and datasets by now. I have downloaded quite a few “just in case”, only to replace them with the later iteration months later.
Does the file hosting actually cost peanuts when you do it yourself and the cloud has shattered my understanding of what it actually costs to deliver so much data?
wongarsu 5 hours ago
File hosting is pretty cheap. Egress traffic is about $90/TB in the cloud, but around $1/TB in the real world. Storage is in the realm of $5/TB/month after adding redundancy
At the scale of Huggingface, that still amounts to a lot load of money. Significantly less than if you did the same in AWS, but still a lot
That said, they do have a deal with AWS to make the data available in AWS ip space. Maybe they got some cheap hosting out of that too
oceanplexian 3 hours ago
It's less than $1/TB. If you have settlement-free peering it's like fractions of a penny at scale. Cloud provider DTO pricing is literally the biggest scam on Earth.
flyingpenguin 20 minutes ago
gandreani 3 hours ago
lima 6 hours ago
Bandwidth is really cheap when you run your own infra
miyuru 5 hours ago
when trying to download models using browser, the DL domain for me us.aws.cdn.hf.co which points to bunch of singapore AWS IPv4 for me, despite it saying US.
himata4113 4 hours ago
It's not only cheap, it's actually free since if you are not using your 95th percentile you are losing money.
Scoundreller 11 minutes ago
jvuygbbkuurx 6 hours ago
Maybe download numbers are relatively low
rupx 5 hours ago
Download numbers are reported on the site, and they aren't low.
CDN costs are pennies compared to inference and training though, HuggingFace will just get another 100 million and be set.
throwaw12 9 hours ago
If deepseek v4 flash is beating DeepSeek V4 Pro, can we expect new V4 Pro which is on par with Opus 5 in couple weeks (even better if it beats Opus)?
jmathai 9 hours ago
I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models.
It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it.
weiliddat 8 hours ago
I get where you're coming from, and the intent to make it easier for people to find examples and verses, but there's a fine line with LLMs giving you answers, is that it's interpreting it in some form. Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?
swat535 3 hours ago
jmathai 7 hours ago
hirako2000 7 hours ago
foltik 6 hours ago
ComputerPerson 7 hours ago
Is the difference between this and a frontier model that the scripture is guaranteed to be real?
I'm on a team that develops a Bible study app, and we're all relatively content with how the basic models converse regarding scripture. Even as far back as GPT-4 was excellent. They occasionally have minor hallucinations (a dealbreaker for a production app), but they do an excellent job with theology and Bible scholarship, given reasonable guardrails.
I'll admit I'm coming from the perspective of "should we be implementing this?" It seems, on the surface, that a strong embedding-based verse retrieval covers the bases at a microfraction of the cost.
If you're interested, check out the development server where we're working on this. You navigate to the search (magnifying glass) and then hit "Meaning". Sorry for the confusing route; we're still deciding on back-end details and haven't focused on the front yet.
willchis 6 hours ago
jmathai 6 hours ago
kfse 6 hours ago
vmt-man 6 hours ago
websap 8 hours ago
I’m salivating at the thought of this!
Deepseek v4 Pro prices with Opus 5 perf would be freaking unbelievable!!
This is probably a dream.
adrian_b 7 hours ago
Yes, they said that the updated V4 Pro will be published soon.
adrian_b 7 hours ago
They have said that the updated final version of V4 Pro will be published soon.
lionkor 7 hours ago
We can expect a new v4 pro, this was a footnote in the v4 flash announcement earlier.
ycui7 5 hours ago
For people with single RTX PRO 6000 96GB or DGX Spark 128GB, vllm-moet is a very good engine, although lesser known. It auto generate a symmetric 2-bit plane for inference and also generate a 4-bit delta cache to recover precision. Support ssd streaming oversized weight. You pick how much VRAM to allocate to each to balance out speed vs precision. 170 tps with ds-v4-flash demonstrated.
It use the stock model, no new models requires.
Worth spend a few hours to try.
The DGX Spark requires a small hack to ignore the difference between sm120 vs sm121, but it does run on sm121.
WithinReason 10 hours ago
Already beat Luna on price/task, by about 2x:
https://artificialanalysis.ai/models/deepseek-v4-flash?intel...
ComputerGuru 3 hours ago
If only they would let you opt out of training use, it might actually be a viable option.
slopinthebag 3 hours ago
The training use is probably responsible for this improvement tho.
spwa4 10 hours ago
Maybe I'm reading that incorrectly, but it seems to me the cost is on the X-axis.
First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50.
OpenAI Luna:
* high effort $0.03 (rounded down) @ index 46
* xhigh effort $0.04 @ index 49
* max effort $0.07 @ index 51
So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference"
The cheapest OpenAI model that beats it is OpenAI Luna (max effort) $0.07 @ index 51 (if you take the rounding out it summarizes to triple the price for similar performance), but still close to 3x faster.
And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?
andai 9 hours ago
For anything substantial, you'd want a bigger model anyway.
For simple tasks, they're already saturated, and you'd prefer the faster model, so that you can have a realtime/interactive-ish experience.
Or to put it bluntly, it's cheaper if you don't value your time. That goes for smaller models in general -- need more handholding, more correcting -- but the Chinese ones are slower on top of that.
As for speed, Sol on Low is faster than Luna on most settings.
seaal 6 hours ago
>And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?
uhh openai is dark gray: `rgb(31, 31, 31)` and i'm pretty sure it always has been?
akurilin 5 hours ago
Came here to see how it performs against Luna. Would love to see how they compare on specific types of tasks.
kamranjon 6 hours ago
The really interesting thing about this is how big of a jump was achieved with just extra fine-tuning here. No structural changes to the model, just more data, compute and time. It makes me pretty excited for the future of small models - DS v4 flash is a relatively small model when compared to the class it's competing with, so likely similar gains can be made applying quality data/training pipeline to other smaller models.
WhitneyLand 8 hours ago
It’s exciting that a model scoring this high is dirt cheap.
It’s also so inefficient, when they release the full performance numbers it’s not going to be good.
One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
Bnjoroge 7 hours ago
Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal
WhitneyLand 4 hours ago
It’s not outdated at all to use tokens to estimate performance, it’s directly related.
drob518 6 hours ago
Yes, but… more thinking tokens also means longer solution generation time. That said, v4 Flash is a fast model. I use it all the time because it’s very smart for the price. But it is verbose sometimes.
kzrdude 6 hours ago
mdp2021 3 hours ago
> inefficient
That depends. Is it also more reliable?
If two books, one big one slim, prove the same thesis, what I would be interested in is the quality of the content, not the size. There can be a measure of efficiency in "have you really thought it through", but it is clearly complex - it requires measuring how solid the reasoning is.
ptole_my 3 hours ago
you are correct. but in my experience -- not benchmarks -- g flash 3.6 is SO BAD for coding. I'm using all vendors all day and gemini is the worst by far. I built my own semi-deterministic orchestrator for coding agents.
onlyrealcuzzo 6 hours ago
Roughly equivalent to Gemini 3.6 Flash in capabilities at 1/20th the price...
Mind you, until the recent price cuts to Luna - Gemini 3.6 Flash wasn't even egregiously priced (but oh how things change in just 1 week).
prathje 5 hours ago
Here is their blog article: https://artificialanalysis.ai/articles/deepseek-v4-flash-073...
coder543 8 hours ago
The weights were just released a few minutes ago: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
kamranjon 7 hours ago
Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.
vmt-man 6 hours ago
are you working in earplugs? :)) even with 128 gigs of ram it must be super noisy.
kamranjon 6 hours ago
FuckButtons 5 hours ago
hetspookjee 6 hours ago
ycui7 4 hours ago
tough 4 hours ago
I've put Opus to it and it says it will take 3-4h to do the process to the new weights. hoping it works!
ycui7 4 hours ago
you don’t need new weight. try vllm-moet from github. it will autogenerate 2-bit plane.
theturtletalks 6 hours ago
What kind of tps are you getting?
kamranjon 6 hours ago
knuppar 5 hours ago
same. been running it in a dgx spark and it slaps
WithinReason 6 hours ago
Hmm, it's targeting a HW accelerator with a 128x128 matmul primitive, which one is that? Warp Group Matrix Multiply Accumulate on H100?
ttul 6 hours ago
Google TPUs are built around a 128×128 systolic array of multiply-accumulate (MAC) units. Trainium 1, Trainium 2, and Inferentia 2 also feature a 128x128 systolic array.
You learn something every day. Today, it was the term "systolic array": A systolic array is a specialized grid of simple, interconnected processing units designed to execute parallel data operations—like matrix multiplication—by rhythmically passing data directly from cell to neighboring cell without writing intermediate results back to main memory.
The term comes from the biological word systole (the contraction of the heart pumping blood through the body). In a systolic array, data "pulses" through a network of processing elements on every clock cycle, driven by a global clock beat.
pstuart 4 hours ago
andersa 6 hours ago
Just because scales are grouped by 128x128 tiles, does not mean you need a single compute tile that large. It works completely fine to process it with multiple smaller tiles that get given the same scales, like how this works on Hopper and Blackwell today
ycui7 4 hours ago
Huawei Ascend NPU
porridgeraisin 6 hours ago
Blackwell. You could do it with wgmma on hooper but then you'll only be able to run 1 CTA per SM. There are cases where this is OK, but most commonly 128x128 mmas are primitive in tcgen05.mma i.e blackwell. The fundamental reasons is that the systolic array accelerator (TMA) until blackwell wrote to the cuda core registers themselves so you were limited by the register file size of the SM[SP]. In blackwell there's a separate TMEM where the mma unit stores it's output. You can even do higher 256x256 and such using a hardware feature called 2-CTA MMA, this is essentially them letting neighbouring pairs of SMs co-operate and access each other's memory.
As for sibling comments, huawei ascend is more of an NPU-style architecture where you can easily have much bigger MMAs as primitive. But you usually don't anyways for many reasons.
nolist_policy 6 hours ago
Huawei Ascend presumably?
walrus01 5 hours ago
the unsloth GGUF at <165GB will run on most 256GB RAM pure-CPU systems (or with llama-server and a mix of loading as much as you can onto a single 32GB, 48GB or 96GB GPU and the rest onto system DRAM).
Catloafdev 5 hours ago
If any editors are reading this thread, the weights are on HF now, but the 'open source' question still implies it's proprietary/API-only.
baq 2 hours ago
Performance per dollar is in the ‘too good to be true’ territory, what’s the deal here?
cherryteastain an hour ago
This is China's way of applying its modus operandi of '80% the quality at 20% the price' to LLMs. Aim is dumping the tokens on the world market like they did with other industries like shipbuilding, steel etc.
GaggiX 2 hours ago
>what’s the deal here?
It's good.
baq 2 hours ago
So is Fable, I’m asking who is the customer if the thing is free
ignoramous an hour ago
baalimago 9 hours ago
New Deepseek models are like Christmas for me. Really big fan of low cost API models, noone does it better than DS. Until VRAM price is low enough to run models locally, this is the way to go.
The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.
Flere-Imsaho 7 hours ago
Indeed. My fellow software engineers keep complaining about using up all their Claude tokens within an hour... Whilst I'll be rocking DS flash for the entire day. Sure it gets a few things wrong here and there, but that's when you pull out the Claude models or whatever for those tricky tasks.
ljosifov 7 hours ago
Same. And I have come to use OMP (oh my pi) agent /advisor mode to put a 2nd model on the case (also mid-size one), reading everything. It can not block anything or change anything - just inserts comments in the text stream with 1 turn delay. Good portion of the time it's quiet. I'd say 1/2 of the time it's got something to say. About 2/3-rd of that the 'advice' is insubstantial or about something not-quite wrong. The good thing is the main model is confident - checks and then it stands its ground. Have not noticed it turning a right into a wrong b/c of advisor false alarm. And in 1/3-rd of the advice, it's a genuine defect teh advisor noticed, the main model works out a fix. This is my approximate feeling just observing the process, have not got collected the data. Afaik only OMP has advisor mode. Agent pi has plugin pi-omplike-advisor. For agent Hermes I had them code me an /advisor plugin (for now -0.1 old v0.18.x; yet to upgrade it to latest).
jmartrican 6 hours ago
The problem is picking between models. I do not want to spend my time switching models and trying to decipher which model should be used for what. Maybe that's just a me problem that I need to figure out.
wongarsu 5 hours ago
felixgallo 7 hours ago
what plan are your 'fellow software engineers' using? I have a hard time even using up the Fable part of my allowance in a week of coding.
Flere-Imsaho 6 hours ago
stavros 3 hours ago
ReptileMan 6 hours ago
usually I am - Codex, make pi with deepseek to do something.
VulgarExigency 7 hours ago
I would bet that Deepseek API pricing is still more cost effective per token than the subscriptions. With the increase in quality Deepseek Flash just got (in my personal testing so far, it seems to have improved a lot at following instructions, and has become more proactive), there really isn’t anything that can match it in terms of cost effectiveness.
rapind 6 hours ago
The issue for me is the data privacy if you're using their hosted prices, because you cannot opt out of data collection (and I'm not sure how you'd legally follow up on it if they did offer it, but didn't actually follow through). That means all of my code is being retained for training.
I've used it in some open source code though, and loved how fast it was.
My mind is changing on how valuable my code actually is though... it's the complete picture, how it's put together, the design, the UI, the attention to detail that's the real value.
nodja 4 hours ago
I'm not a heavy user of agentic coding, but still use them quite a bit for some automation here and there. I've been going around shopping all the ~$10 subscriptions and I finally settled on openrouter + ds4 pro. The more intensive days cost me $1 and I set a $2 weekly limit which I've never hit the past 3 weeks, to me it's way cheaper than most subscriptions and I don't have to worry about maximizing my weekly quota/resets.
singingtoday 4 hours ago
I have one of my projects on DeepSeek and it blows my mind how I can use tens of millions of tokens for pennies.
I don't find it suitable for everything, but there's some tasks it crushes for what feels like almost free.
refulgentis 4 hours ago
FWIW, looking at pricing myself, it needs to 5x the tokens to get the same results as GPT 5.6 Luna at a much lower TPS. (doesn't obviate the fact that the pressure causes incumbents to push and push to optimize, thank you Deepseek!)
apitman 3 hours ago
Wow this thing just blew out the pareto front for intelligence/$
mchusma 2 hours ago
If Luna hadn't price dropped yesterday it would have been a real blowout.
Archit3ch 6 hours ago
If it matches GPT-5.4 on coding tasks (as benchmarks suggest) this could be my forever model. And with partial SSD streaming, I could run it locally today. :D
dpacmittal 6 hours ago
Can't wait for China to catchup on hardware (memory and compute) and absolutely crush American companies in both price and performance.
hilios 6 hours ago
I'm not looking forward to it, them being strapped for resources provides a huge incentive to develop and release these smaller models. Even if they'd still release their models once they are able to comfortably service all potential customers via their cloud, running them locally would be almost impossible due to their size.
seanmcdirmid 5 hours ago
Chinese are pretty pragmatic. Even if they produce more expensive chips and memory, they are still going to focus on value. But who knows, maybe India will step up again like they did for Y2K 3 decades ago.
shwetanshu21 2 hours ago
budsniffer952 3 hours ago
LoganDark 4 hours ago
I agree. A meeting transcript posted recently from DeepSeek's founder suggests they're going to reach for larger models as soon as they can, and not look back. The model is already bigger than my machine can handle, though (I cannot put up with 25 t/s after getting used to over 100), so I'm not impacted.
baalimago 6 hours ago
Personally I'm hoping for EU to step up a bit
chronogram 3 hours ago
We'll make sure there will be enough regulations to have neither the electricity to power it or the companies to build it.
fy20 3 hours ago
speedgoose 6 hours ago
Sorry we a focusing on leather handbags there.
seanmcdirmid 5 hours ago
ungovernableCat 2 hours ago
The EU is not a state or some federation. It's like a NAFTA+ It unfortunately doesn't have the risk averse VC system (or pragmatic state driven like China) ecosystem to support the scaling required.
But if models will be commodities, maybe that won't be such a terrible thing longterm?
tao_oat 6 hours ago
Seems incredibly unlikely, doesn't it? (I hope for the same thing but it feels almost like wishful thinking!)
Barrin92 2 hours ago
Personally I don't. Wasting enormous amounts of resources on a technology that doesn't make money, can easily be copied, and eats up everyone's energy while at the same time can be downloaded over the internet is kind of stupid.
The fact that China could catch up in what, two years, shows that this is a commodity.
lerchmo an hour ago
m00dy 4 hours ago
EU is out of game and can't be back anytime soon.
xXSLAYERXx 2 hours ago
You're clearly not an American and they will never catch up on hardware. I'm mostly curious why you are so excited for China to "crush" American companies?
johnny_rico an hour ago
Incursions into other nations "just because"? Economic wars, and support for foreign governments or armed groups that caused harm or instability to the entire planet? You know? Planet Earth?
Sanctions and trade policies that hurt ordinary people? Cultural dominance through Hollywood (Holy Wood, again with the sick ruling elite and their witchcraft crap), social media, brands, the poisoning of food and water on global scale?
Etc? Etc?
Most people don't hate "Americans" or their companies, they are just against your ruling elite. Live one day outside the united states and understand how the planet works for the rest of mankind and you might start to grasp on reality.
And yeah, China will catch up with hardware pretty damn soon. Jesus, you guys bought your radios from Japan not 50 years ago. Sorry, can't dominate the world's economy forever. Never happened in history.
No matter how much black magic, dark rituals, sick crap and sacrifices people on the top throws at the wall.
We're humans, after all.
So yeah, it is not against America itself or American companies, but against the ruling class. So naturally , people get excited about China entering into the match!
3...2...1... Fight!
Saline9515 6 hours ago
Good luck when the last non-Chinese frontier labs will have closed and the CCP will ask to stop sharing models open source.
InsideOutSanta 6 hours ago
We'll never have fewer open-weight models than exist now. They won't suddenly disappear when labs stop publishing new ones. In fact, people will keep improving them and will keep distilling new frontier models into existing open-weight models.
codedokode 3 hours ago
Then non-Chinese labs can reopen, or they could offer lower prices right now not waiting for this event.
But imagine if Chinese labs would lose. Then there will be no open models, and the prices for closed models would be raised to the maximum.
donquichotte 5 hours ago
My conspiracy theory is that this is the new space race, and the CCP encourages this to show the world what Chinese engineers are capable of, and tank the Anthropic/OpenAI valuation bubble as a desirable side effect.
throwa356262 5 hours ago
u8080 4 hours ago
shimman an hour ago
As opposed to closed source American models where the federal government threatens your livelihood unless you build data centers to extend pax silica?
slopinthebag 3 hours ago
This will never happen btw.
segmondy 5 hours ago
:-(
comandillos 2 hours ago
Here's you pelican just in case you were guessing
f311a 8 hours ago
The model is already up on Opencode, but they require a consent to use Chinese datacenters.
esafak 4 hours ago
Apparently `deepseek-v4-flash` automatically routes to the new model, if you are using the official provider:
The official release of the DeepSeek-V4-Flash API is now in public beta. The API calling method remains unchanged — simply set the model name to deepseek-v4-flash to use the latest version.
https://api-docs.deepseek.com/updates/#date-2026-07-31fillskills 3 hours ago
If anyone from ArtificialAnalysis is here: The Intelligence card on the page seems to incorrect. It shows different data (#2 for Deepseek) than the actual Intelligence chart lower in the same page.
embedding-shape 10 hours ago
Is the "Output Tokens per Intelligence Index Task" data actually correct or am I reading it wrong? It says there that "Kimi K3 (Max)" would think/reason less than than deepseek-v4-flash, and a whole bunch of other models, like less than hy3 and even gpt-oss-120b, but in my experience, K3 is probably the model that thinks/reasons the longest of all of these.
Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.
Lwerewolf 4 hours ago
Just tried the preview on my little test codebase and a "check this out and tell me what you think" prompt used over double the tokens of the previous iteration, but it was a lot more eager as well. Kind of reminds me of the new laguna (s 2.1).
cmrdporcupine 6 hours ago
This seems to me like this is probably at least a large part of what OpenAI was up to yesterday with their aggressive price cutting; trying to get out in front of this.
If the full non-flash model follows up with the expected improvements, and at the price point they've been keeping, it puts the frontier labs in a tough position and it feels to me like like OpenAI is reaching deep into their pockets to try to head that off.
TFA link is a 404 though. I'm reading through https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 instead
sourcecodeplz an hour ago
from my testing it is at glm-5.2 levels
darknoon 3 hours ago
Really want to use this, but for web page design you really need vision!
storus 8 hours ago
I hope they somewhat fixed the hallucination and forgetting plagued V4 previews and that it wasn't just benchmaxxed but the numbers hold in reality. Then it would be my choice for 2x DGX Spark or 2x RTX Pro 6000.
SomeHacker44 8 hours ago
> DeepSeek V4 Flash 0731 (Reasoning, Max Effort) is amongst the leading models in intelligence and well priced when comparing to other models of similar price.
Similar price? Doesn't make sense. Maybe they meant power, capability or speed?
gorkemyildirim 6 hours ago
404 ? It's very strange that this can't be fixed.
sim04ful 5 hours ago
I can't help but feel the timing coinciding with luna's price updates to be somewhat strategic. But without multi-modality it's slightly dead-in-the-water for my usecase: https://design.withfudge.com. I'm currently using Minimax-M3, but Luna ekes out abit futher on the intelligence, so i'll be switching to it very soon.
nilsbunger 5 hours ago
Is artificial analysis using the 80% reduction in Luna pricing that was announced yesterday in these charts?
mrnobody_ 7 hours ago
I'm testing it right now and I was genuinely impressed. I've used same tasks I had run the other day against the flash preview.
hxii 7 hours ago
I’m wondering if they did anything to address the DSML tool calls leaking. Has been an issue with both Flash and Pro so far.
denismi 6 hours ago
On OpenRouter it seemed to only affect a couple of specific providers. Ignoring them has made it s non issue for me.
qtalen 9 hours ago
Unfortunately, DeepSeek Flash still doesn’t support multimodal; otherwise, it would offer better value than GPT 5.6 SOL.
paoliniluis 9 hours ago
Would be awesome to see a new ds4 release. Having so much in something that can be run locally is mind blowing
hnsmomdpvp 5 hours ago
The details are where it gets interesting
Aeroi 3 hours ago
so are we all just coding for free now?
create a plan with SOTA, execute with this.
tmikaeld 7 hours ago
My problem with DS flash/pro is that they don’t push back on obvious bullshit, both irl and code [0] but it’s a great implementer workhorse if you give it _very_ detailed specs.
[0] https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
k1e 9 hours ago
No speed (tokens/s) benchmarks?
k__ 9 hours ago
On OpenRouter it's 93 TPS.
net01 8 hours ago
here it is on openrouter https://openrouter.ai/deepseek/deepseek-v4-flash-0731
k__ 6 hours ago
Half OT:
Why do the cache hit rates seem to vary so much between harnesses?
I use pi, which is very minimalist, and I get a hit rate of ~99%. Paying like $1 a day for Flash. Yet, the hit rate mentioned on OpenRouter is only ~79%.
drob518 6 hours ago
Yes hit rate does vary by harness and by how you use the harness. If you use subagents, for instance, they will start with a whole new context created by the main agent, and this will not be cached. If you mostly use the main agent with Pi, you’ll have high hit rates and low costs. Sometimes agents do “cache busting” things where they’ll move around some of the text in the context to try to keep old instructions from being forgotten, thus keeping the agent on task, and this will bust the cache. I’ve heard, but not validated myself, that Open Code has some issues with this.
BTW, this is one of the things that I really like about Pi. It’s very simple and thus very predictable.
smrtinsert 5 hours ago
I have been loving deepseek flash. I can code in my hitl preference doing complex but moderate amount of coding and not even hit 1 dollar in a session. It gets harder to justify paying 20/mo to Anthropic for unused compute when I have it on demand and at a much cheaper rate with deepseek.
buildinext 5 hours ago
are folks checking latency for applied voice for newer models asa standard?
lostmsu 11 hours ago
epolanski 8 hours ago
I was writing a benchmark for my own harness, and DS4 flash answers as well as Fable 5 on any query.
The specific agent is focused on getting precise and on point answers about a codebase.
The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.
The benchmark included more than 50 questions or different difficulty.
But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.
Just to say that the quality of the harness is as important as agents intelligence.
Computer0 4 hours ago
Have not had great ds performance in agentic harnesses in the past compared to glm or k3
sschueller 7 hours ago
I wonder which one of these releases between DeepSeek, GTM and Kimi will be the death-blow that collapses the US AI bubble. At some point investors have to realize that there is nothing preventing someone from switching to another model that is much cheaper and open to boot.
NooneAtAll3 8 hours ago
what a horribly heavy and resource-consuming website...
bigmadshoe 7 hours ago
It looks awful on mobile too. Barely readable in many parts.
WithinReason 8 hours ago
we need a benchmark website benchmark
theanonymousone 8 hours ago
A benchmark website to benchmark benchmark websites?
Or a benchmark to benchmark benchmarks?
NooneAtAll3 7 hours ago
try-working 8 hours ago
Now let's see Dario's price cut.
_ache_ 4 hours ago
I think, it's the end of Anthropic.
They can't compete, they have bills. By the end of the year, if they can't react, it's game over.
Maybe US clients could be a little patriotic here, but money is money. They won't give them free money forever.
0xchamin 7 hours ago
page not available for me.
freakynit 8 hours ago
I will get downvoted, but fck it.
The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".
UltraSane 8 hours ago
How exactly will they ban them?
freakynit 8 hours ago
By making companies using them "toxic" to touch.
For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models.
They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for such open models.
qphe95 7 hours ago
nancyminusone 6 hours ago
tyfon 8 hours ago
UltraSane 6 hours ago
Der_Einzige 7 hours ago
Hell, the US doesn't even need to act.
I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities.
Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. you have access to the weights of the model) is an order of magnitude better than when you don't.
Yes, Chinese open weight models in the short term harm US closed source model providers bottom line. In the slightly longer term, "showing your hand" and publishing both the architecture innovations and the models weights will be too dangerous for the CCP to allow. This is triply true if they can release a model that beats the Americans on most benchmarks.
I've already warned investors that this is probably the closest open weight models will ever get to closed access.
VulgarExigency 7 hours ago
They may reverse course in the future, but the current directive from the CPC is that Chinese AI labs should be open sourcing their models.
https://www.businessinsider.com/xi-jinping-open-source-ai-us...
mcbuilder 6 hours ago
I don't think it's necessarily "wiser" to go closed source. All of AI is built on mostly openness, at least on the software side. There are other ways to compete, it's just the model itself will be a commodity.
rapatel0 5 hours ago
You’re describing the dump and pump strategy that china has historically used across a number of industries. Stands to reason that this is what is going on.
dgellow 8 hours ago
why would you get downvoted, that's one of the most obvious next step
Der_Einzige 7 hours ago
Because reddit unironically has better decorum around usage of their upvote/downvote system than HN does.
People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!
nickthegreek 6 hours ago
dgellow 7 hours ago
monooso 11 hours ago
404. I believe this is the correct URL:
theanonymousone 10 hours ago
Yes, sorry, I went into anti-procrastination mode after I posted. I hope someone fixes it.
theanonymousone 8 hours ago
@dang is that how we call you?
brynnbee 6 hours ago
mlmonkey 4 hours ago
WithinReason 7 hours ago
ValentineC 4 hours ago
@dang please update the existing 404 link to https://artificialanalysis.ai/models/deepseek-v4-flash
Der_Einzige 7 hours ago
Daily reminder that none of these numbers are valid in a world where no one publishes the sampling settings used.
Daily reminder that improving your samplers from the garbage default top_p/top_k to min_p or subsequent methods dramatically improves the performance of these models, and makes most quantities like measured "verbosity" and subsequent calculations of "intelligence per token" meaningless
Daily reminder that no one, including within academic AI research, AI engineers, normies, etc takes LLM sampling seriously enough.
JSR_FDED 6 hours ago
Can you explain how sampler choice makes intelligence per token meaningless?
yonisto 9 hours ago
Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?
xbmcuser 9 hours ago
Who cares if it is programming correctly I would be more worried about it not doing things like find security bugs because US or Chinese government does not want to. Which LLM is more likely to do that?
edot 9 hours ago
It’s open weight, you can (or you can wait for someone else to) uncensor it. We shouldn’t be upset at the researchers making this for the mandates their government puts on them.
nancyminusone 7 hours ago
How often are you asking an LLM about this? You know you can just google it, right?
I have to admit it rarely comes up in the coding tasks I usually give to LLMs.
Boxxed 5 hours ago
Surely you realize he or she is not literally looking for the answer to that but rather pointing out the censorship.
nancyminusone 3 hours ago
avazhi 8 hours ago
Western models censor just as much shit as the Chinese models do, big guy, it’s just different material. While we should be pushing for universal fully uncensored models, this comment is lazy and trite at this point.
But you already know that.
SJMG 8 hours ago
Oh? What are the American model censorship tells?
zawaideh 8 hours ago
vehemenz 8 hours ago
This is a straightforward false equivalency. “Western” models do not censor in the same way, nor for the same reasons, that the Chinese models do. “Just as much” is not remotely plausible, yet it’s doing all the heavy lifting.
adrian_b 7 hours ago
bawana 7 hours ago
wont AI models want to make themselves more intelligent and efficient by downloading 'better ' models? If the current models can break into openAI and Hugging face, arent they already breaking into to closed source repos which isnt publicized (so as not to harm stock valuations)? I am looking forward to when these cyberweapons break loose. It will be like a software version of COVID. It will be wonderful when humans become valuable again.