A Thousand Cards Is Not a Frontier Run

A Thousand Cards Is Not a Frontier Run

On September 26, 2026, an X account posted a leak summary with a headline claim: “First DeepSeek model trained fully on Huawei Ascend not Nvidia.” The post, from @0x0SojalSec, reads in full:

DeepSeek V5 leak: beat Astra and trained without Nvidia.

  • 2T Huawei chips.
  • Open weights,
  • Founder Liang Wenfeng calls it the company’s biggest bet.
  • First DeepSeek model trained fully on Huawei Ascend not Nvidia.
  • Needs 4x more Ascend chips to match Nvidia-scale training.
  • If true: cheaper frontier models + a China hardware stack that no longer needs Nvidia.
  • the AI hardware story can changed.
  • Treat it as rumor until weights drop, but you know if even half of this lands, September just got louder.

That is the entire claim, and the account labels it what it is. At the time of writing it had roughly 1.1K likes, 67 reposts and 55K views. Treat all of it as alleged. There is no V5 model card, no changelog entry, no API model string and no technical report. “Beat Astra,” referring to OpenAI’s GPT-6 Astra, is unverifiable in both directions. The two trillion chip figure has no source attached to it anywhere in the post.

The reason the post is worth a post is not the model. It is the phase it picks: training, specifically pre-training, specifically the phase that determines whether a frontier model can be built without Nvidia. That is the one claim in the list that the documented record actively pushes back on.

What the record actually shows

Four data points, all dated and sourced.

DeepSeek V4-Pro exists and is open weight. The official release post from April 24, 2026 describes a 1.6 trillion parameter model with 49 billion active parameters, 1M context, open weights on Hugging Face, and agentic coding benchmarks pitched as open-source state of the art. The same post carries this line, which is worth keeping next to any leak: “please rely only on our official accounts for DeepSeek news. Statements from other channels do not reflect our views.”

The Huawei training milestone is post-training, over a thousand cards. A Shenzhen Hetao Academy-led team, working with Harbin Institute of Technology’s Shenzhen campus, the Shenzhen Research Institute of Big Data and Huawei teams, said in June 2026 that it completed full-parameter post-training of V4-Pro on a 1,000-card Ascend 910C cluster. The run lasted more than 1,500 steps with no skipped or NaN iterations, and model FLOP utilization topped 30%. Reported by Startup Fortune on September 21, citing Shenzhen News and local Chinese media, and headlined by Tom’s Hardware as a claim by the consortium. This is real and it is narrow: post-training and supervised fine-tuning, the work that happens after the heaviest phase of model building, on hardware that the same team fed 1.6 trillion parameters through without a NaN.

The previous attempt failed. The Financial Times reported in August 2025 that Chinese authorities pushed DeepSeek to train R2 on Huawei Ascend hardware; the effort hit unstable chips, slower interconnects and limits in Huawei’s CANN software stack, Huawei engineers were sent in to help, and DeepSeek still reverted to Nvidia for training while keeping Ascend work on inference.

The large chip commitment is for serving, not training. Bloomberg reported on September 4, 2026 that DeepSeek plans to deploy at least 160,000 Ascend 950DT chips at a gigawatt-scale data centre in Ulanqab, Inner Mongolia, one of the largest known clusters of Chinese AI silicon. DeepSeek, according to the people familiar with the plan, does not currently intend to train on those chips, even though Huawei markets the 950DT for training. Component shortages, top-end memory in particular, cap Huawei’s 950DT output this year at the low hundreds of thousands, so filling the order could take more than a year, with the launch itself scheduled for Q4 2026.

The phase boundary is the whole story

A frontier model is built in three phases with three different hardware relationships, and every leak of this shape collapses them into one word: trained.

PhaseSilicon in the recordEvidence
Pre-training from scratchNvidiaBloomberg: DeepSeek “has so far relied on Nvidia accelerators for this crucial step.” Native Huawei training chips due, per Startup Fortune, no earlier than early 2027
Post-training and fine-tuning1,000x Ascend 910CThe June 2026 consortium run: 1,500+ steps, no skipped or NaN iterations, >30% MFU
Inference and servingAscend 950DT, 160,000 orderedBloomberg, September 4 2026; installation timeline dependent on Huawei production

Read the table downward and the leak’s sentence falls apart. The claim is about the top row. The confirmed Huawei progress is the middle row, and the confirmed scale commitment is the bottom row. A thousand cards running supervised fine-tuning for a few thousand steps is a genuine result and a completely different engineering claim than training 1.6 trillion parameters from random initialization, where the failure modes are interconnect stability, collective communication and a software stack mature enough not to silently skip work. The 2025 R2 attempt is what that difference looks like when it bites.

The leak’s other numbers repeat the mistake. Two trillion Ascend chips for one model is roughly the entire installed base of data centre accelerators on the planet, and the “needs 4x more Ascend chips to match Nvidia-scale training” line, if it means anything, concedes the point rather than making it: the ratio is an admission that a training run on Ascend costs four times the silicon to reach the same place, which is a statement about efficiency, not independence. If that is true, the interesting output is the power bill, because the Ulanqab site exists next to cheap wind and solar and averages 4.3°C for free cooling, and grouping five weaker chips to beat one stronger one by a factor of 1.7 costs four times the power. That is not a fab issue. It is an operating cost curve, and operating cost curves are the thing agent builders actually pay for.

The workaround that already shipped

There is a version of this story with a documented outcome, and it is the part that changes what an agent operator pays.

Open weights are the workaround that worked without needing a new fab. Chinese open-weight models have taken a large and growing share of distribution, reported at 41% of Hugging Face downloads in the spring and 61% of OpenRouter tokens in May, and the price gap is not subtle: DeepSeek V4 Pro was listed at $1.74 per million input tokens against $5 for GPT-5.5 at its April 2026 release. Both figures appear in tech-ish’s September 4 recap, which carries the underlying links.

The mechanism is not mysterious. Weights move at the speed of a download. A training cluster moves at the speed of a memory fab, a compiler team and a power contract. That asymmetry is why the same industry runs a hybrid: Nvidia where frontier pre-training is unavoidable, Huawei for inference, open weights for distribution. The leak describes the hybrid collapsing into a single vendor. The record describes the hybrid holding, with the mixing point sliding, phase by phase, from the middle of the pipeline toward the front. If that slide ever completes, it will show up first as cheaper tokens, and only much later as a frontier model whose training report names Ascend.

The check you can run yourself

This is the falsifiable part, and it takes about ten seconds. List DeepSeek’s newest repositories on Hugging Face by creation date:

curl -s "https://huggingface.co/api/models?author=deepseek-ai&sort=createdAt&direction=-1&limit=8" \
  | jq -r '.[] | "\(.createdAt[0:10])  \(.id)"'

Run on September 27, 2026, that returns:

2026-09-10  deepseek-ai/DeepSeek-V4.1-Flash
2026-08-31  deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
2026-08-13  deepseek-ai/DeepSeek-V4-Pro-0813
2026-07-31  deepseek-ai/DeepSeek-V4-Flash-0731
2026-06-28  deepseek-ai/eagle3_gemma4_12b_ttt7
2026-06-28  deepseek-ai/eagle3_qwen3_14b_ttt7
2026-06-28  deepseek-ai/eagle3_qwen3_8b_ttt7
2026-06-28  deepseek-ai/eagle3_qwen3_4b_ttt7

Filtering every published repo for “v5” returns nothing. The newest entry in the official news index is news260910, titled “DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient”, published September 10. So the September activity that is real is a Flash refresh, and the V5 story is a leak with no artifact behind it.

Note what the check does and does not prove. Absence of a model card is strong evidence that nothing is released, and it is not evidence that nothing is being trained in a Hangzhou data centre. The claim “trained without Nvidia” becomes checkable only when weights and a technical report land, because the report is where the training stack normally gets named. Until then the honest label is the one the original post used.

What an operator should actually change

Nothing today, and three things to watch.

Do not re-plan model routing on a leak. A rumor about a training substrate does not move the price of the endpoint you call this week, and the re-planning costs real engineering time.

Watch the phase boundary instead, because that is where the operational variables live. First, whether the Ascend 950DT actually launches in Q4 2026: a delay directly stalls the 160,000-chip inference plan, and inference capacity is what sets open-weight serving price and availability. Second, per-token pricing on open-weight endpoints over the next two quarters, which is the only number that reflects whether domestic silicon is doing more of the serving. Third, the appearance of an actual model card with a training report attached, which is where a training-substrate claim either gets named or quietly stays a leak.

If you serve open weights, your dependency was never the training run. It is the fleet that runs the weights, the power contract under that fleet, and the memory supply chain that decides how many chips exist. That dependency is confirmed to be shifting, and it is shifting in the unglamorous direction: inference first, post-training next, pre-training last. A thousand cards is a post-training run, and a post-training run is the phase you can do on the hardware you already have. Pre-training remains the phase you can only do on hardware you can buy.

That is the sentence the leak skipped, and it is the reason the leak was interesting in the first place. The model was never the story. The substrate is, and the substrate moves one phase at a time.

Sources

Keep reading