The market saw 106% revenue growth and $21.34 billion in free cash flow. I saw a three-word anomaly buried in the guidance. 73.5% to 74.5%. Q3 gross margin guidance, one full point below Q2's actuals. In a quarter where NVIDIA holds over 90% of the AI training GPU market and sells every chip it can fab, margins don't dip by accident.
Tracing the gas leaks before the code compiles.
The leak is in the packaging line. CoWoS-L. Two compute dies. Eight HBM3e stacks. A 2.5D interposer pushing more bandwidth than most data centers had five years ago. And a yield curve that hasn't caught up to the demand curve.
The Numbers Behind the Numbers
NVIDIA's FY2025 Q2, ending July 28, 2024, was a blowout by any conventional measure. Revenue grew 106% year-over-year. Adjusted gross margin hit 74.5%. Operating cash flow came in around $25 billion. Free cash flow: $21.34 billion. The company bought back roughly $5 billion in stock and paid dividends. All of this happened while hyperscalers โ Microsoft, Google, Amazon, Meta โ funneled over $200 billion in combined AI capital expenditures toward NVIDIA's order book.
The headline story is simple: AI compute demand is real, structural, and nowhere near saturation. Channel inventory sits under 30 days. H100 and H200 units sell for $25,000 to $40,000 per chip. The upcoming B200 is expected to price at $50,000 to $70,000. NVIDIA isn't selling chips. It's printing allocation tickets.
But I didn't spend four months auditing Golem's ICO contract in 2017 to read headline numbers. I learned to read the debug logs. And the debug log here is the margin guidance.
The CoWoS Constraint
Here's the technical reality. NVIDIA is a fabless designer. It doesn't own fabs. Its entire AI chip output depends on TSMC's advanced process nodes โ 4N for Hopper, 4NP for Blackwell โ and, more critically, on TSMC's CoWoS advanced packaging capacity. CoWoS utilization is running near 100%. TSMC is doubling capacity through 2024, but demand still outstrips supply.
The packaging math matters more than the transistor math. H100 uses CoWoS-S. B200 moves to CoWoS-L, a larger interposer with higher interconnect density. CoWoS-L supports two reticle-sized compute dies and eight HBM3e stacks. That's a step function in complexity. And complexity, in semiconductor manufacturing, translates directly to yield risk.
Industry reports put early B200 yields in the 60-70% range. Hopper's 4N process is mature, running above 90% yield. Blackwell's 4NP is new. The gap between those numbers is the gap between Q2's 74.5% gross margin and Q3's guided 73.5-74.5%. Initial yield ramp costs money. Packaging defects get scrapped. Every failed CoWoS-L interposer is a $50,000 chip that never ships.
This is the hidden signal in the earnings report. NVIDIA didn't call it out explicitly. It didn't have to. The guidance did the talking. Silence between the blocks tells the real story.
The Prepayment Strategy
There's a second signal hiding in the cash flow statement. NVIDIA's free cash flow of $21.34 billion sits below its net income. In a normal company, that gap raises questions. For NVIDIA, it's the answer to a different question: how do you secure supply when your suppliers are bottlenecked?
NVIDIA has been paying large prepayments to TSMC and SK Hynix to lock CoWoS capacity and HBM3e allocation. This is "hidden capital expenditure" โ the company keeps its own capex light at 5-8% of revenue, but the actual capital commitment flows through prepayments on the balance sheet. It's a smart move. It's also a signal. NVIDIA is so confident in demand that it's willing to tie up billions in prepayments to guarantee future supply.
The alternative โ waiting for spot market capacity โ doesn't exist. CoWoS is sold out. HBM3e from SK Hynix, Samsung, and Micron is allocated months in advance. If you're not prepaying, you're not shipping.
The Ecosystem Moat
Now let me address the elephant in the room. AMD's MI300 series. Google's TPU. Amazon's Trainium. Microsoft's Maia. Everyone wants a piece of the AI compute pie.
Here's what the hardware comparison misses: CUDA.
I've written before about how liquidity is just patience with a time limit. The same logic applies to software ecosystems. CUDA is a 15-year accumulation of developer mindshare, libraries, and optimization tricks that no hardware spec sheet can replicate. AMD can match transistor counts. It can match memory bandwidth. It cannot match the fact that every AI researcher, every ML engineer, every data scientist learned to work in CUDA.
The market share numbers confirm this. NVIDIA holds over 90% of the AI training GPU market. About 80% of AI inference. The second-place competitor, AMD, sits at 5-10%. This isn't a competitive race. It's a structural monopoly built on a software moat.
But here's the contrarian angle: the real threat isn't AMD. It's the customers themselves. Hyperscalers represent about 54% of NVIDIA's revenue. Microsoft alone is 15-20%. These customers are developing in-house silicon not because they want to replace NVIDIA, but because they want negotiating leverage and cost optimization for specific inference workloads.
The risk isn't that NVIDIA loses the training market. The risk is that inference โ projected to surpass training demand by 2025 โ becomes a fragmented market where specialized ASICs compete on cost-per-token rather than raw performance. In that world, NVIDIA's 80% inference share faces gradual erosion.
The China Question
There's also the geopolitical overlay. China's share of NVIDIA's revenue has dropped from roughly 20% in 2023 to about 10% in 2024. Export controls on high-performance AI chips โ those with TPP โฅ 4800 โ have effectively locked NVIDIA out of its second-largest market.
The market treats this as a minor headwind. The U.S., Europe, and the Middle East are picking up the slack. Sovereign AI demand from Middle Eastern and Southeast Asian governments is becoming a new growth vector.
The two-year horizon looks fine. The five-year horizon is less certain. China's national AI chip initiative โ backed by a 344 billion yuan third-phase semiconductor fund โ is funding domestic alternatives from Huawei and Cambricon. The hardware gap is closing. The software gap remains, but Chinese developers don't have access to CUDA the way U.S. developers do. They're building alternatives out of necessity.
The rug wasn't pulled overnight. It's being pulled thread by thread, year by year.
What the Margin Dip Actually Means
Let me bring this back to the trade. The Q3 margin guidance of 73.5-74.5% isn't a red flag. It's a tell. It tells you that Blackwell is shipping in volume, that the yield curve is still climbing, and that NVIDIA is willing to accept margin compression in the short term to dominate the next product cycle.
The math is straightforward. If B200 yields improve from 65% to 80% by mid-2025 โ which is the expected trajectory as TSMC matures the 4NP process and CoWoS-L packaging โ gross margins should recover to 75% or higher. The margin dip is a bridge, not a destination.
The real question isn't whether NVIDIA can ship Blackwell. It's whether the AI capex cycle holds. Microsoft, Google, Amazon, and Meta have committed over $200 billion to AI infrastructure. If AI application revenue doesn't materialize at the pace these capex budgets assume, the next earnings cycle will feature a different kind of margin story.
The Signal to Watch
Two weeks in the lab, one second in the field. I've spent years debugging markets and models. The signal I'm watching isn't NVIDIA's gross margin. It's the hyperscaler earnings calls. When Microsoft or Google signals a pause in AI infrastructure spending, that's the exit signal.
Until then, the setup is clear. Blackwell ramps through Q4 2024 and scales through 2025. CoWoS capacity doubles. HBM3e supply improves. NVIDIA's pricing power holds. The margin dip is temporary. The moat is structural.
The market is pricing NVIDIA for perfection. The question isn't whether NVIDIA executes. It's whether the AI demand curve stays steeper than the supply curve. For now, both curves are moving in the same direction. The tape says buy. The code says wait for the next hyperscaler earnings print.
Debug the market. Not the headlines.