DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

DSeek plans to officially release the V4.1 Flash model around September 10, 2026 (Beijing Time). After extensive internal and external testing, V4.1 Flash has comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time. In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support!

We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.

121 points | by nickweb 2 hours ago

15 comments

  • jiehong 36 minutes ago
    Sounds nice!

    But, the web ui chat version of flash has very poor language following abilities in my experience:

    You may ask it something in English, and get a thinking chain in Chinese with an answer in Chinese, or an English thinking chain and an English answer. Using the retry button on the same question has a 50/50 chance of any of those results.

    Sometimes, asking something in English, but where information are mostly in another language may make the answer in the language where data has been found. The other day, I asked something about a local German thing, in English, and I got an answer in German instead. It’s as if all the language data stirred it away from the language of the user’s question.

    • apexalpha 15 minutes ago
      I have the same issue, sometimes.

      I initially thought it was a trick, that using Chinese chars is somehow more info dense and it saves tokens to 'think' in Chinese.

      But later on it became more erratic. I still wonder if token reduction would work that way.

    • el_io 27 minutes ago
      I'm also totally not sure why it do that, but I guess because they're searching from China and web results comeback in Chinese so the model start using that.
    • swiftcoder 34 minutes ago
      I've hit this too, but you can just add "in English" to steer it
    • epolanski 26 minutes ago
      I've occasionally got chinese characters in anthropic/openai's responses too, locally on codex/claude.

      Hasn't happened in a while, last time was when I was testing fable 5 in june.

      • oefrha 4 minutes ago
        I don’t know what model codex uses for session summarization (I use Pro subscription, no third party models), but I get Chinese summaries from time to time, when the only Chinese that could have appeared in the session would be an i18n strings file that it may or may not have loaded. Very puzzling. Last happened yesterday.
  • oefrha 1 hour ago
    Source is apparently a banner announcement on https://platform.deepseek.com/usage. Had me searching for a couple minutes...
    • nickweb 1 hour ago
      I swear I put that at the start of the post. Must've managed to miss it when copy and pasting!
  • postalcoder 47 minutes ago
    I hope DeepSeek takes some time to improve their tuning for reasoning effort. Right now, there are only three reasoning efforts: low, high, and max.

    For all intents and purposes, "low" is pretty much the same as turning reasoning off, and "high" is similar to "max". "High/max" performs way too much reasoning, takes forever, and causes costs to balloon. They need a proper "medium" setting.

    I get it that they're probably focused on pushing performance right now, but the ergonomics of the model aren't great.

    • tarruda 23 minutes ago
      I would rather have just 3 levels: low, medium and high.
    • iamniels 28 minutes ago
      I switched to GLM-5.3 flash on high for this reason. Too many "but wait" in the Deepseek-v4 reasoning.
  • EbNar 48 minutes ago
    Since a few months, I almost exclusively use the Chinese "flash" models for my needs. They are a joy and they cost pennies per answer. Great job.
    • pimeys 15 minutes ago
      Yes. I'm working in the agent industry and my god are we excited on new versions of Chinese flash models. The direct competition is Gemini Flash, and these models are much better on agentic tasks with fraction of the task price compared to Gemini. Things like oh here's a set of simple instructions for you to follow, call these tools, return this report. 70-80% of the price. And especially Deepseek Flash produces better quality than Gemini does.

      Where Gemini still wins is non-text input what Deepseek cannot do, yet, and Deepseek Flash has this thing of cheaper models where a failing tool call can derail your agent to a retry loop if you're not careful on instructions in the error message.

      If they fix and make the tool calls to work better in non-optimal situations, it's much easier to switch from Gemini without a few weeks of evals and bugfixing.

    • serf 6 minutes ago
      I recently had to config my harness to watch for cybersecurity flags from astra and funnel requests to flash when they occur because Astra gets queezy when you talk to it about UDP packets in games.

      Works fantastic. Glad there is a more 'uncensored' thing to fall back to when the frontier folk are too sensitive.

    • ActionHank 37 minutes ago
      I am legitimately more excited for this release than any frontier models at this point.

      I don't need a model that can invent new mathematics. I need something that is fast, cheap, and consistent. Give me that and I can build and scale.

    • darkoob12 31 minutes ago
      My mental bias always kept me away from Chinese models. Because i know that china is a surveillance state and all the things we know about CCP. But after what we learned about OpenAI and how they most likely used user data to basically cheat in an open competition i think it does not matter which AI provider you use all of them will own your data and all of them can spy on you. So I am willing to switch to Chinese models. This way we help them develop and improve models some day we can run them locally.
      • Mashimo 19 minutes ago
        The new meta model is fast and very cheap as well, and when used through OpenCode you get quite a lot of free tokens. But meta is also THE surveillance company, so probably also not a good choice in your case.
      • ricardobeat 29 minutes ago
        These models are open-weights. Anyone can host them, you don’t have to use chinese servers even though most of them offer zero data-retention policies.
      • el_io 24 minutes ago
        You can use those models from Openrouter, they have many Non-Chinese providers.
      • epolanski 25 minutes ago
        You're naive if you're thinking the scumbags running the US companies aren't using your data.

        In any case old rules apply: if privacy is a concern don't share the data. I share all my work-related code because it's worthless, but I don't and would never share company business and process details, access to production/user data, etc.

        Meanwhile I know of people connecting all the kind of MCPs for datadog/sentry/jira/concluce/production databases to their harnessess..lol.

    • XzAeRosho 35 minutes ago
      Same for me. DeepSeek models are incredibly good at implementation and light planning. I still default to Opus models for feature planning, but for most simple features the Pro models suffice.

      Incredible good value and product they have built.

  • tarruda 1 hour ago
    Hopefully it will be open weights and have the same architecture and size as the current v4 flash vision, which is probably the best LLM that can be run on 128G devices.
  • hinow 26 minutes ago
    But what no one mentions is that the price is going from a starting point of $0.16 to $0.60, so basically they're charging nearly four times as much.
    • _aavaa_ 24 minutes ago
      What are you talking about? Current flash prices are 0.66 for output, this is dropping it to 0.60.
  • tensegrist 25 minutes ago

        In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price.
    
    just in terms of user perception when selling this sort of service, this is what they call a "good look"
  • nickweb 1 hour ago
    Via nitter: https://xcancel.com/JustinGorya/status/2097287080128708930

    Looks like the new model can be used if summoned via the API but the API won't list it.

  • indigodaddy 12 minutes ago
    So, will it have vision? (based on deepseek-v4-flash-vision-exp ?)
  • igleria 1 hour ago
    v4 pro was decent then a better cheaper faster model comes now?

    As a consumer I feel like hansel and gretel combined, deepseek could be the witch.

    • calgoo 41 minutes ago
      v4 flash has been working quite well for the majority of my personal projects, with occasional v4 pro or Kimi 3 for the most complicated tasks or to check the overall project progress (when vibe coding).
      • HansHamster 17 minutes ago
        I must be doing something wrong. I gave v4 pro a try a couple of days ago, gave it a simple prompt like "clean up functions x and y in file z" and it would always start off promising, just to quickly get sidetracked, start hallucinating problems in the code, and just get stuck for hours until I interrupt it:

        — hmm — 0x2D696370 — little-endian bytes: 70 63 69 2D = 'p','c','i','-' — hmm — WAIT — WAIT — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *HOLD ON — HOLD ON — HOLD ON — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — !!!!!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *OK — WAIT — I THINK I FINALLY SEE THE WHOLE PICTURE — I NEVER READ IT — AND — THE LAYOUT — hmm — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — HOLD ON — HOLD ON — HOLD ON — HOLD ON

        Then gave the same to Sonnet 5 and it was done 15 - 30 minutes later. I tried v4 pro both in claude code and codewhale with similar results. Haven't tried the new deepseek harness.

    • throwaway473825 51 minutes ago
      It's not unprecedented given that GLM 5.3 Flash was better and cheaper than GLM 5.2.
  • swiftcoder 1 hour ago
    If they can keep up this cadence of Flash leap-frogging the previous Pro, we're in for a good time
  • nicce 54 minutes ago
    > In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support!

    Wow. Imagine OpenAI/Google/Anthropic doing this! Nope.

  • thrownaway561 51 minutes ago
    I will continue to be amazed by how much power you get from DeepSeek Flash for the cost. I have let that puppy lose on so many projects and it is has never let me down. It can build and entire Rails app in no time and even do the tests. For most things, I don't get why people pay the money for Claude. DeepSeek Flash is my default agent in Omarchy.
  • neugls 2 hours ago
    Waiting to use it
  • donk8r 53 minutes ago
    [flagged]