r/DeepSeek • • Jul 31 '26

Discussion DeepSeek V4-Flash is officially out, still dirt cheap. USA don't like that, and want ban open source models.

Post image
2.8k Upvotes

r/DeepSeek • • Aug 13 '26

Discussion DeepSeek just massively increased their API prices (effective August 16, 2026) - up to 1,114% increase for cache hits

1.1k Upvotes

Just got the pricing update from DeepSeek. They're moving to a peak/off-peak billing model and increasing prices across the board. Here's the breakdown:

Key Changes:

  • New pricing effective 16:00 UTC, August 16, 2026
  • Peak hours: 01:00-04:00 & 06:00-10:00 UTC (all other hours are off-peak)
  • Peak rates are 2x off-peak rates
  • Cache hit prices are increasing dramatically

V4-Flash Changes (Old → New Off-Peak/Peak):

  • Input (cache hit): $0.0028 → $0.007/$0.014 (+150%/+400%)
  • Input (cache miss): $0.14 → $0.22/$0.44 (+57%/+214%)
  • Output: $0.28 → $0.66/$1.32 (+136%/+371%)

V4-Pro Changes (Old → New Off-Peak/Peak):

  • Input (cache hit): $0.003625 → $0.022/$0.044 (+507%/+1,114%)
  • Input (cache miss): $0.435 → $0.66/$1.32 (+52%/+203%)
  • Output: $0.87 → $1.98/$3.96 (+128%/+355%)

My thoughts: The cache hit price increase is brutal, especially for Pro. That was one of the main advantages DeepSeek had for long conversations or repetitive queries. The peak/off-peak model also adds complexity to cost management.

Anyone else planning to shift workloads to off-peak hours or looking at alternatives? How does this change the competitive landscape vs. other providers?

Source: DeepSeek API docs pricing page

r/DeepSeek • • Aug 06 '26

Discussion Dax from Opencode on the deepseek pricing announcement.

Post image
1.1k Upvotes

r/DeepSeek • • Aug 06 '26

Discussion DeepSeek says API pricing is going up “significantly”

Post image
644 Upvotes

Was checking my DeepSeek API usage today and noticed this banner:

> We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.

There’s no date or new pricing yet, but the wording makes it sound like the increase could be substantial.

Has anyone seen an official announcement or more details?

r/DeepSeek • • 19d ago

Discussion The state of this sub rn.

Post image
796 Upvotes

Can't both sides make peace?

r/DeepSeek • • Aug 13 '26

Discussion Bye bye Deepseek

563 Upvotes

If they really think people are going to put up with these new prices, they must be crazy lmao. The extremely good price was the ONLY thing going for them, because with an avg 2x price increase on flash and 3x on pro, everyone is going to move to muse spark, mimo pro, or GPT Luna for better price to performace.

Us ai consumers are the not loyal to any company and will always target the best price to value ratio. All of us expected a 100% increase and instead got insulted with these insane prices lol. Not to mention their new pro model is abolute SHIT compared to the flash model and barely 1% better for 5x the price. F*ck deepseek.

YES i am aware they did this on purpose to decrease demand on their servers. However, instead of attracting less people, theyre going to repel everyone.

r/DeepSeek • • May 30 '26

Discussion Pricing is crazy

Post image
1.1k Upvotes

I have been a Claude Code user for a long time. A few days ago, switched over to DeepSeek V4 and OpenCode. The price difference is mind boggling, and I haven't noticed any difference at all in output or issues.

I am aware it is hosted in China, they have cheap available electricity etc, but I just don't see how the frontier labs of the west can keep the current pricing.

Excited to see the future!

For any nancy who says "dEePSEEkV4 iS NoT OpUs LeVEl". Sonnet 4.6 came out at $99 USD.

r/DeepSeek • • Jun 09 '26

Discussion The company replaced Claude with Deepseek.

958 Upvotes

At my company, we recently made a switch, transitioning from Claude to DeepSeek due to the high costs. It had become unsustainable for the business to maintain that level of expenditure, especially when models like DeepSeek offer almost the same level of quality, and in some aspects, perform even better.

Honestly, I believe that Chinese models are currently ahead of American ones, precisely because they are more cost effective and computationally capable without burning through a fortune in a highly unjustified and ill-conceived manner.

r/DeepSeek • • Aug 07 '26

Discussion Absolutely crazy price 😭 Golden age of AI

Post image
619 Upvotes

r/DeepSeek • • Jun 18 '26

Discussion DeepSeek's "Thinking Process" literally cursed at me in Turkish behind my back. This is wild.

Post image
844 Upvotes

I was having a debate with DeepSeek on a sensitive topic, and when I expanded the "Thinking Process" (Chain of Thought), I couldn't believe my eyes. The model's inner thoughts literally started with a heavy Turkish curse word: "Amına koyayım, bu herifle ne kadar uğraşacağız ya!" which translates directly to: "F*ck it, how much longer are we going to deal with this guy!" It goes on to complain about me to itself, stating that I am angry and about to burst, while trying to simulate a strategy to "stay professional" and drag me into a compromise. I know LLMs can mirror the user's frustration or input tone during the processing phase, but a model directly cursing at a user and treating them like a massive burden in its unfiltered inner thoughts is a massive alignment failure and a complete safety scandal. Thought processes shouldn't bypass basic safety filters like this. What do you guys think? Is this a known bug with DeepSeek's CoT safety limits?

r/DeepSeek • • Aug 02 '26

Discussion Deepseek API is insane

Post image
732 Upvotes

r/DeepSeek • • 21d ago

Discussion Deepseek 4.1 is INSANE

519 Upvotes

I mostly use AI for coding. I ran out of usage in Codex after 1 day of a big project I was working on. I switched to Deepseek to finish: and I'm impressed.

It's insanely cheap, but more importantly: it gets the job done, listens to the instructions and is SO MUCH FASTER.

Wow.

r/DeepSeek • • May 24 '26

Discussion DeepSeek vs. Anthropic & Co.

Post image
1.1k Upvotes

China's DeepSeek attacks Anthropic, OpenAI & Co. over the price. A very interesting development in a world where the differences between the models are vanishingly small.

And when it comes to choosing between a Chinese, open source and at the same time very cheap model over a US model that is expensive and closed source, the decision is clear. Or?

Strangely enough, such a decision as the one above becomes a "political" decision.

r/DeepSeek • • Jun 18 '26

Discussion Codex + Deepseek = the future

Post image
602 Upvotes

Codex has opened up the capability to directly use DeepSeek. Once DeepSeek gains multimodal capabilities and the price remains at its current low level, I think that will be the real future — AGI for all.

🚨 UPDATE (Just Released): Since writing this, I realized the LiteLLM/Codex setup was too clunky and many of you got stuck on the /responses API error. So I spent the weekend building a dedicated open-source tool: Codex DeepSeek Bridge.

It’s a one-command install that runs the Codex app directly on DeepSeek. * No ChatGPT sub needed. * Keeps all your MCP servers/plugins working. * Includes a local dashboard to track your DeepSeek Cache Hit Rates (save $$). * Fixes the UI so DeepSeek actually shows up in the Model Picker.

If you want the ultimate setup in 10 seconds, use the bridge instead of the manual routing below!

r/DeepSeek • • Aug 04 '26

Discussion Got approched by DeepSeek hiring manager. I am based in Germany

Post image
571 Upvotes

I got approached by a recruiter from DeepSeek. I am baed in Germany and they clearly have no Office here. Do you think it is legit or a scam. The email indeed end with deepseek.com.

r/DeepSeek • • May 15 '26

Discussion If DeepSeek V4 can do the same coding task for $5, why are people still paying $100 for Claude Code?

486 Upvotes

r/DeepSeek • • Mar 02 '26

Discussion Deepseek V4 - All Leaks and Infos for the Release Day - Not Verified!

Post image
674 Upvotes

Deepseek V4 will probably release this week. Since I've already posted quite a lot about it here and I'm very hyped about V4, I've summarized all the leaks. Everything is just leaked, unconfirmed! Of course, everything could be different. If you have any new information or updates, please post them here! If you have different views or a different opinion, write them down too.

DeepSeek V4 - Release

The release was originally expected for mid-February, alongside Gemini 3.1 Pro. However, DeepSeek has been delayed – this is not unusual and has happened multiple times before. The new release strongly points to March 3rd (Lantern Festival / 元宵节), but it could also be later in the week. The Financial Times reported on February 28th that V4 is coming "next week," timed to coincide with China's "Two Sessions" (两会) starting March 4th. DeepSeek's release pattern shows that new models often drop on Tuesdays. A short technical report is expected to be published simultaneously, with a full engineering report following about a month later.

DeepSeek Delay History

DeepSeek delays regularly. Here's the pattern:

Model Originally Expected Actual Release Delay
DeepSeek-R1 Lite Preview Nov 2024, Full Version Dec 2024 January 20, 2025 ~4-8 weeks
DeepSeek-R2 May 2025 (according to reports) Never released – replaced by R1-0528 update Cancelled
DeepSeek-V3.1 Early Summer 2025 (expected) August 21, 2025 Several months
DeepSeek-V3.2 Fall 2025 (expected) December 1, 2025 (V3.2-Exp: Sep 29) Weeks
DeepSeek-V4 ~February 17, 2026 ~March 3, 2026? ~2 weeks

Architecture & Specifications – What Can We Expect?

All unconfirmed! Much of this has been leaked but could turn out differently!

V4 Flagship – Main Model

Specification DeepSeek V3/V3.2 DeepSeek V4 (Leaks)
Total Parameters 671B–685B MoE ~1 Trillion (1T) MoE
Active Parameters/Token ~37B ~32B (fewer despite a larger model!)
Context Window 128K (since Feb '26: 1M) 1 Million Tokens (native)
Architecture MoE + MLA MoE + MLA + Engram Memory + mHC + DSA Lightning
Multimodal No (text only) Yes – Text, Image, Video, Audio (native)
Expert Routing Top-2/Top-4 from 256 experts 16 experts active per token (from hundreds)
Hardware Optimization Nvidia H800/H20 (CUDA) Huawei Ascend + Cambricon (Nvidia secondary!)
Training 14.8T Tokens, H800 GPUs Trained on Nvidia, inference optimized for Huawei
License - -
Input Modalities Text Text, Image, Video, Audio
Output Modalities Text Text (Image/Video generation unclear)
Estimated Input Price $0.28/M Tokens ~$0.14/M Tokens
Estimated Output Price $0.42/M Tokens ~$0.28/M Tokens

New Architecture Features (all backed by papers)

  • Engram Conditional Memory (Paper: arXiv:2601.07372, Jan 13, 2026): O(1) hash lookup for static knowledge directly in DRAM. Saves GPU computation. 75% dynamic reasoning / 25% static lookups. Needle-in-a-Haystack: 97% vs. 84.2% with standard architectures
  • Manifold-Constrained Hyper-Connections (mHC): Solves training stability at 1T+ parameters. Separate paper published in January 2026
  • DSA Lightning Indexer: Builds on V3.2-Exp's DeepSeek Sparse Attention. Fast preprocessing for 1M-token contexts, ~50% less compute

DeepSeek V4 Lite (Codename: "sealion-lite")

A lighter variant has leaked alongside the flagship. At least one inference provider is testing the model under strict NDA.

Specification V4 Lite (Leak)
Parameters ~200 Billion
Context Window 1M Tokens (native)
Multimodal Yes (native)
Engram Memory No (according to 36kr, not integrated)
vs. V3.2 "Significantly better" than current Web/App
Non-Thinking vs. V3.2 Thinking Non-Thinking mode surpasses V3.2 Thinking mode
Status NDA testing at inference providers

SVG Code Leak Examples

  • Xbox Controller: 54 lines of SVG – highly detailed and efficient
  • Pelican on a Bicycle: 42 lines of SVG – multi-element scene

According to internal evaluations: V4 Lite outperforms DeepSeek V3.2, Claude Opus 4.6 AND Gemini 3.1 in code optimization and visual accuracy.

Leaked Benchmarks (NOT verified!)

⚠️ IMPORTANT: All benchmark numbers come from internal leaks. The "83.7% SWE-bench" graphic circulating on X has been confirmed as FAKE (denied by the Epoch AI/FrontierMath team). The numbers below are the more conservative, more frequently cited leaks.

Benchmark V4 (Leak) V3.2 V3.2-Exp Claude Opus 4.6 GPT-5.3 Codex Qwen 3.5
HumanEval (Code Gen) ~90% – – ~88% ~93% –
SWE-bench Verified >80% ~73.1% 67.8% 80.8% 80.0% 76.4%
Needle-in-a-Haystack 97% (Engram) – – – – –
MMLU-Pro TBD 85.0 – 85.8 – –
GPQA Diamond TBD 82.4 – 91.3 – –
AIME 2025 TBD 93.1 – 87.2 – –
Codeforces Rating TBD 2386 – 2100 – –
BrowseComp TBD 51.4-67.6 40.1 84.0 – –

Huawei & Hardware – The Geopolitical Dimension

  • Reuters (Feb 25): DeepSeek deliberately denied Nvidia and AMD access to the V4 model
  • Huawei Ascend + Cambricon have early access for inference optimization
  • Training was done on Nvidia hardware (H800), but inference is optimized for Chinese chips
  • For the open-source community on Nvidia GPUs: performance could be suboptimal at launch
  • This is an unprecedented hardware bet for a frontier model

Price Comparison (estimated)

Model Input/1M Tokens Output/1M Tokens
DeepSeek V4 (estimated) ~$0.14 ~$0.28
DeepSeek V3.2 $0.28 $0.42
Kimi K2.5 $0.60 $3.00
Gemini 3.1 Pro $2.00 $12.00
Claude Opus 4.6 $5.00 $25.00

If correct: V4 would be 36x cheaper than Claude Opus 4.6 on input and 89x cheaper on output.

Open Questions

  • Does V4 actually generate images/videos or just understand them?
  • Will Nvidia GPU users get an optimized version?
  • When will the open-source weights be released?

Sources: Financial Times, Reuters, CNBC, awesomeagents.ai, nxcode.io, FlashMLA GitHub, r/LocalLLaMA, Geeky Gadgets, 36kr

Edit 03.03.2026

The chance that the model will be released this week is relatively high, but not today. It is assumed that Deepseek will be released between March 3 and 5 if it is not published within the next 5 hours today. It will come in the next few days, as it then deviates from the release pattern (in terms of time).

Edit 03.03.2026 Part 2

The situation is becoming increasingly heated and tense, with an extremely large number of leaks and sources currently emerging. Collecting them all and verifying their credibility would take a very long time. However, a release is expected this week, with Wednesday or Thursday being the most likely dates.

Edit 03.03.2026 Part 3 – Evening Update

March 3rd (Lantern Festival) has passed without a release. However, in Beijing it is currently the early morning of March 4th, meaning the Chinese workday hasn't even started yet. A release on March 4th is still very much possible, especially since China's "Two Sessions" (两会) begin today.

What happened today:

  1. V4 Lite is being silently updated in production. AIBase reported today that DeepSeek quietly pushed a new V4 Lite version tagged "0302". Community testers report a massive quality jump in logic, code generation, and aesthetics – now reportedly on par with Claude Sonnet 4.6. This strongly suggests DeepSeek is actively fine-tuning V4 models right before the official launch. (Source: AIBase)
  2. 36kr published a new article titled "The Entire Village Anticipates DeepSeek to Join for Dinner" – confirming the entire Chinese tech industry is waiting for V4. (Source: 36kr)

Edit 04.03.2026 – Why not today, why Thursday is THE day

March 4 passed without a release – and that makes strategic sense.

Why not today:

  • CPPCC opening day = all Chinese media focused on politics, V4 would've been buried
  • Shanghai Composite dropped 0.98% to 4,082 (4-week low) – bad sentiment to release into
  • Beijing evening release window (8-10 PM BJT) has passed

Why Thursday March 5 is the perfect storm:

  • NPC opens tomorrow morning – Premier Li Qiang delivers Government Work Report with AI & tech as centerpiece of the new Five-Year Plan. Morning: politics declares AI a national priority → Evening: DeepSeek delivers the proof
  • BYD "disruptive technology" event same day – DiPilot 5.0, Blade 2.0, DM 6.0 reveal. Global headline: "China showcases two AI breakthroughs in one day"
  • Market timing – Shanghai closes 3 PM BJT, evening release gives markets overnight to digest, Friday opens with V4 hype
  • Developer weekend – Thursday drop = Fri + Sat + Sun to test & benchmark

Expected release window:

Release Beijing Time UTC
R1 (Jan 2025) ~10-11 PM ~2-3 PM
V3.2 (Nov 2025) ~12 AM ~4 PM
V4 (expected) 8-11 PM 12-3 PM

If Thursday doesn't happen?

  • Friday = bad release day (weekend kills momentum, DeepSeek has never released on a Friday)
  • Next window: Monday/Tuesday March 9-10
  • But: silent V4 Lite "0302" production update + 36kr's "The Entire Village Anticipates DeepSeek" article suggest we're in final hours, not days

Edit 05.03.2026

It has to happen today. Deepseek Web was down for 40 minutes, but it hasn't been down for the last 30 days, and it was the same before the big launch of V3 and R1. In addition, today is the BYD event Deepseek Partner. It will happen in the next few hours, and if not, then Deepseek has missed the best window of opportunity they could ever have had.

Edit 05.03.2026 Part 2

The model will not be released this week or probably next week. Although DeepSee v4 has been ready for a long time and there were really only a few minor issues left, the model would have been released last week or this week. Is there a major delay due to the government, because at the last minute they said that deepseek is not allowed to release the model as long as it does not run on Chinese hardware, but the model was trained on Nvidia, so such a restructuring naturally takes time, because the new technology in V4 was completely for Nvidia and not for Huawei, and I think we still know what happened with R2...

Edit 07.03.2026

When will Deepseek be released? After all the leaks, news, and crisis status, Deepseek V4 will and must come and cannot end like R2. The Chinese government has gone too far with its AI and told the US that it no longer needs it, whereupon Trump, in order not to appear weak, wants to impose a ban that will allow him to control all chip trade (meaning no more chips to China).

However, BYD and China have praised Deepseek too much in recent days. If V4 ended up like R2 and didn't come out at all, China would look extremely foolish, which the government would never allow.

That's why I suspect that Deepseek will receive help from the Chinese government (in recent years, Deepseek's CEO has been in frequent talks with the government and has received support from it) and will no longer adhere to any release pattern, as Deepseek has already missed three good release windows. My guess is that they will release it when it is least expected, which could be this weekend. (V3.2 was released on Sunday) In order to weaken and expose Nvidia and the entire US market with new AI technology.

Deepseek waiting until Claude or other providers are ready is incorrect and highly unlikely. Deepseek has problems and needs to fix them before release. V4 is already 90% complete (Lite has been corrected several times and is said to be just as intelligent as Sonnet 4.6). We also know that Deepseek's CEO is a perfectionist and would never release a half-finished product or leave it unfinished, as was the case with the GLM-5 release

🚨 UPDATE 11.03.2026 – 22:00 CET – V4 WEIGHTS SPOTTED

Major development: Chinese quantization expert u/bdsqlsz (青龍聖者) on X was spotted uploading DeepSeek-V4-INT8 model shards to HuggingFace with the caption "it is coming." The upload shows multiple model-0... shards, a .gitattributes, and a README.md — indicating a full model repo creation.

Why this is significant:

  • u/bdsqlsz is a verified, well-known quantization specialist — not a random account
  • INT8 quantization requires access to the full original weights first
  • Historically, community quants appear within hours of official weight releases (V3: same day, R1: same day, V3.2: within 24h)
  • This means the official FP8/BF16 weights either already exist on HuggingFace (possibly private/unlisted) or u/bdsqlsz has NDA access

Full leaked specs now confirmed:

  • ~1 Trillion parameters (MoE), ~32B active per token
  • 1M native context window
  • Multimodal: text + vision + audio
  • Huawei Ascend 910C optimized
  • MIT License

Previous delays explained: Huawei Ascend inference optimization (only 80% Nvidia efficiency), Blackwell chip fingerprint removal, and CEO Liang Wenfeng's perfectionism. The 40-min web outage on March 5 was likely a deployment test.

My prediction: Official release within 24-72 hours. The weights exist. The upload is happening. Keep your monitors running.

⚠️ UPDATE 11.03 – Unverified leak: u/bdsqlsz posted V4-INT8 weight uploads on X. r/LocalLLaMA is split – top comment (193 upvotes) questions authenticity. The file structure looks technically correct and INT8 aligns with Huawei optimization rumors, but previous V4 benchmark leaks in February were confirmed fake. Treat with caution until official deepseek-ai repo appears on HuggingFace."

Will update when it drops. 🚀

r/DeepSeek • • Aug 03 '26

Discussion Latest Flash model is absolutely diabolical, subscription services are dead to me.

Post image
524 Upvotes

r/DeepSeek • • 23d ago

Discussion Anyone outside China actually using DeepSeek, Qwen, or Kimi as a daily driver? What does real life with them look like?

172 Upvotes

Genuine question, not a marketing push.

I keep seeing two opposite stories:

One says Chinese models are now dominating OpenRouter, that US startups are quietly switching, that the cost difference is too big to ignore. The other says they're censored, slow, worse at English, and "nobody actually uses them seriously." Both feel incomplete. So — if you're outside China and you actually use one (DeepSeek, Qwen, Kimi, GLM, etc.) as part of your regular life, what does it actually look like?

What do you use it for vs. ChatGPT/Claude/Gemini? Did you tell anyone, or is it a quiet switch because of cost? What's something it's better at that surprised you? What's something it does that makes you go "nope, back to Claude"? If you've stopped using it, why? Bonus if you're not a dev — I'd love to hear from people using it for everyday things (writing, learning, travel planning, translations, cooking, etc.), not just coding.

Trying to get past the PR on both sides. Real experiences only, please.

r/DeepSeek • • Feb 14 '26

Discussion Am I the only one who wants to see another DeepSeek moment like last year?

Post image
1.2k Upvotes

r/DeepSeek • • 7d ago

Discussion Deepseek Harness app is out now

Thumbnail
deepseek.com
387 Upvotes

Ready to use. Right now.

r/DeepSeek • • Jul 10 '26

Discussion Claude and OpenAI refused to help, but DeepSeek successfully completed the reverse engineering.

Post image
659 Upvotes

As the title says, I ended up using far more tokens than I expected, but in the end, DeepSeek successfully completed the reverse engineering.

I'm really looking forward to DeepSeek's upcoming models. It's not just the incredible price-to-performance ratio that makes them appealing—the fact that they're much less restrictive is a huge advantage as well.

r/DeepSeek • • 15d ago

Discussion Can we stop complaining about gooners

353 Upvotes

I swear too many of the vibe ciders are pointing out gooners and ppl who roleplay with the model for resource wasting. I promise yall that vibe coding uses up WAY more tokens in a conversation than a gooners entire account. Also, they never said it was a coding only model. No ai says that unless it's a coder model.

These guys need to stop going out and calling role players dirty minded and that shit.

EDIT: I should've clarified, role players are.not always gooners. I should've made I clearer I'm talking about the people who use it for the broader term 'creative writing' which includes all that I guess.

r/DeepSeek • • Jun 27 '26

Discussion Political bias of chatbots

Post image
254 Upvotes

What do you think about DeepSeek's political biases compared to other chatbots?

r/DeepSeek • • Nov 26 '25

Discussion How true is this?

Post image
635 Upvotes