China is cooked...
Skills:
Staying Current in AI80%
Key Takeaways
The video discusses the release of NVIDIA's Groq 3 LPU and its impact on the AI industry, particularly on Chinese labs, highlighting the significant cost efficiency and throughput gains of the new chip compared to Chinese A100/H100 chips.
Full Transcript
Ever since the Deepseek moment in January 2025, the openweights market has been largely split between DeepSeek, Meta, and Mistrial. Now, obviously since then, the market's been pretty saturated where we now have the decline of DeepSeek and the rise of Miniax, Kimmy, OpenAI, and Quen have mixed into the competition. But when we look at Open Router sampling of the market in terms of token usage, closed labs still beat the entire open-source models. And the common rhetoric is well how will closed labs remain competitive long term? Wouldn't the open models catch up eventually? When we talk about competition between open models mostly from China against closed models mostly from companies like OpenAI, Anthropic and Gemini, we have to look at the economics of these companies as much as we do for benchmarks. We have to start looking at access to compute. In my previous video, I mentioned even among NeoClouds, their access to newer graphics cards in itself is a competitive edge. And here's the same thing. I'm here at Nvidia GTC and Nvidia just released their VR Rubin platform. And if you look at the Gro 3 LPU that runs under Gro 3 LPX, Open Labs in China will have a really hard time getting access to compute modules like this. To give perspective on why this is such a competitive edge is because when you use Gro 3 LPU, Nvidia reports that you could get up to 35 times lower cost in token while delivering 50 times throughput per megawatt which is only possible through hardware acceleration that Chinese labs will have a harder time getting access to while American hyperscalers will be the first to access. In other words, access to newer compute like VR Rubin NVL72, their VR CPU racks, Gro 3 LPX and Bluefield 4 STX and Spectrum 6 are all equipment that closirth firstand access to compared to open Chinese labs. Now, in comparison, Chinese labs have been saying how they were computrained for a long time. The head of Quen from Alibaba, Justin Lynn, has reflected how Chinese labs beating US labs in the next 3 to 5 years is less than 20% likely, largely due to lack of compute availability. The LPX rack comes with 256 of these LPU processors which will all add up to 128 GB altogether that serve as SRAM module which can give up to 640 terabytes per second in throughput which in comparison to HPM module that is around 1 TB per second. You can see this insane difference in what you can expect from compute module like this. For clarity, the Gro 3 LPU is not meant to be an entirely new chip with their own CUDA, but rather they're an accelerator where you can essentially offload parts of the work that would benefit from high throughput like the feed forward network layer. Now when we look at having equipments like this even if Chinese labs are able to remain competitive in benchmarks for a long period of time when we look at the metric like 10 times more revenue opportunity and 35 times more throughput per megawatt the demand for closed labs will continue to grow and their operating margin will also likely shrink which allows closed labs more runway to compete. If you have been following my content for some time, I have been tracking the release cycle from labs and how we will see model releases at a faster cycle and equipment like this will change the innovation speed. Even when we start to look at new VR Rubin NVL72, they're offering values like needing 75% less graphics cars to train and 10% of the cost in training. Now that doesn't certainly mean that there will no longer be a place for Chinese labs to continue innovating. But what we will likely going to see the industry converge into only few labs remaining because the compute gap will make labs especially Chinese lab that have chips that are 3 to four years old in technology and given how behind Huawei chips are in comparison more and more Chinese lab will not make sense economically to continue competing in this rather capitalintensive race. Now to extend an olive branch, we still have no real grasp on how big the openclaw market is going to be. In other words, the demand for openclaw might be the rising tide that lifts all boats. Meaning the rather smaller demand profile we see from token traffic study from open router survey. The overall demand for AI might grow, making the economics of Chinese labs still viable since people will still want to use small to medium-siz models that they can run locally. But unless they're able to monetize heavily from releasing their models open in short of continuing to be relevant and if Nvidia chips like the Gro 3 LPU continue to allow hyperscalers here in the US to deliver each token far cheaper and faster, that means the cost of intelligence will go lower, allowing the Frontier Labs to start offering their input and output tokens for pricing that is competitive to Chinese lab without losing their margin at a higher speed. What do you think? Do you think Chinese lab will be squeezed out of the market given where the industry is heading towards?
Original Description
NVIDIA just released their Groq 3 LPU along with their other vera rubin chips and the gap of compute and token availability has grown even more from frontier labs vs Chinese labs.
The cost efficiency gain and throughput gain on the upcoming Groq 3 LPU chip in comparison to Chinese A100/H100 chips are becoming obviously huge.
#ai #nvidia #china #llm #technology
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: Staying Current in AI
View skill →Related Reads
📰
📰
📰
📰
Stop Showing the World in Midjourney: Make the Viewer Enter It
Medium · AI
Is It Too Late to Learn AI? I Did the Arithmetic on All 8.3 Billion of Us.
Medium · AI
What If Your LLM Could Never Make Up a Fake Source Again?
Dev.to AI
What secretly eats your local LLMs' speed as your context fills up - Part 3
Dev.to AI
🎓
Tutor Explanation
DeepCamp AI