
NVIDIA CEO Says GPT-6 Astra Used 100,000 Blackwell GPUs

NVIDIA CEO Says GPT-6 Astra Used 100,000 Blackwell GPUs
WEEX View
- The key variable is whether NVIDIA’s disclosed cluster scale becomes a new benchmark for frontier model training rather than a one-off build. If it does, access to advanced GPUs and interconnect capacity may become even more concentrated among a small group of top AI companies.
- The report also points to a second watchpoint: generation gap matters as much as raw chip count. Blackwell, Hopper, and domestic alternatives are not directly interchangeable, so comparisons based only on total units may understate the performance gap.
- For crypto-adjacent AI narratives, the practical implication is not immediate token impact but renewed focus on compute scarcity, infrastructure bottlenecks, and whether smaller players can realistically compete without access to top-tier centralized hardware.
NVIDIA CEO Jensen Huang said GPT-6 Astra was trained using more than 100,000 Grace Blackwell GPUs connected in a high-speed NVLink72 cluster, according to the company disclosure cited in the report. Huang also said the next batch would use 400,000 GPUs, though the specific models were not identified.
The disclosure centers on the hardware used to train GPT-6 Astra. Huang said the system used more than 100,000 Grace Blackwell GPUs linked through NVLink72, NVIDIA’s architecture for placing 72 GPUs within the same high-speed interconnect domain. The report said the next training batch is planned at 400,000 GPUs, but did not specify whether those systems would also use Blackwell.
The report framed the disclosure against the current AI infrastructure gap between U.S. and Chinese model developers. ByteDance was described as the closest among major Chinese model companies, with about 36,000 B200 GPUs connected. Its domestic clusters were said to rely mainly on Hopper-based H20 and H800 systems.
Other capacity figures in the report remain less clear. Kimi was said to have obtained around 20,000 Hopper GPUs through Alibaba, specifically H200 units, although Alibaba denied that claim. DeepSeek has not disclosed the full training hardware used for V4, but leaked information cited in the report said the company had about 20,000 H-equivalent compute units in May, including roughly 16,000 units of Huawei 950 capacity.
The report also noted that the Blackwell platform used for Astra is no longer NVIDIA’s newest generation. Vera Rubin has already entered full-scale production, and NVIDIA estimates that training large mixture-of-experts models with Rubin could require only a quarter of the GPUs needed with Blackwell. No Chinese company has publicly disclosed using 100,000 advanced GPUs of the same generation to train a single model, according to the report.
Why It Matters
The disclosure matters because it sharpens the divide between frontier AI development and the broader market narrative around model competition. At the top end, performance is increasingly tied not just to algorithms or data, but to who can assemble, power, and interconnect extremely large clusters of the newest chips.
That has broader implications beyond AI vendors themselves. The more large-model progress depends on scarce, tightly integrated hardware stacks, the more strategic weight shifts toward semiconductor supply, networking architecture, and access to advanced compute. For crypto-linked AI sectors, that keeps the focus on infrastructure constraints rather than simple enthusiasm around AI branding.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
About WEEX View
WEEX View is a crypto analysis and intelligence hub, covering the latest in Web3, AI, and global markets. Get independent research and in-depth insights to stay ahead of market trends and trading opportunities.
Latest articles
MoreLedger CTO Urges Coordinated Disclosure on Wallet Flaws
Ledger CTO Charles Guillemet called on hardware wallet vendors and researchers to follow coordinated vulnerability disclosure, warning that faster AI-assisted flaw discovery could raise risks for users if issues are publicized before fixes are ready.
Solana Raises Transaction Size Limit to 4096 Bytes
Solana has increased the maximum size of a single transaction from 1232 bytes to 4096 bytes, expanding the amount of data and logic developers can fit into one transaction for more complex on-chain applications.
White House Crypto Advisor Backs CLARITY Act Ahead of Senate Vote
White House crypto advisor Patrick Witt said skepticism around the CLARITY Act will be proven wrong as the Senate prepares for a September 15 vote on the digital-asset market structure bill.
Bitcoin-Gold Correlation Reaches Highest Level Since 2017
Bitcoin’s 90-day correlation with gold has climbed to 0.56, the highest level since 2017, while its correlation with the Nasdaq 100 fell to about 0.30, pointing to a shift in cross-asset market behavior.




