FUNDA

FUNDA

Deep|NVDA: Rethinking NVIDIA's Moat in the AI Stack

FUNDA's avatar
FUNDA
Apr 23, 2026
∙ Paid

NVIDIA — The Bellwether of the AI Era

NVIDIA’s hardware leadership is virtually undisputed. In many respects, the company is not only one of the most successful enterprises of the AI era but also one of the principal infrastructure providers enabling this generational shift in compute — a position that we think deserves full respect.

Over the past several years, the certainty of NVIDIA’s roadmap execution has been rare: from Blackwell to Vera Rubin, and onward to Rubin Ultra and Feynman, the cadence has held essentially at one generation per year. At the same time, NVIDIA has not confined itself to the GPU. It has continued to extend at the system level — whether by folding inference capability (such as the LPU architecture) into the platform, pushing co-packaged optics to pave the way for hyperscale clusters, or reshaping the boundary between storage and GPU memory through directions like Storage-Next. At heart, these moves reinforce its “AI factory”-grade system capability. Combined with deep entrenchment in advanced packaging capacity, NVIDIA has erected exceptionally strong moats along both the hardware and the supply-chain dimensions.

However, focusing solely on hardware leadership would still understate NVIDIA’s core competitiveness. What we think deserves closer attention is its profit structure: Data Center already accounts for the overwhelming majority of revenue, the chip-level gross margin for the GPU is roughly 84%, and overall GAAP gross margin reached 75% (FY2026 Q4). In the semiconductor industry, this is almost an outlier — closer to the profitability profile of a software company than that of a conventional hardware vendor.

Precisely for that reason, we believe NVIDIA’s current success is driven not by hardware alone but by its end-to-end pricing power across systems, software, and ecosystem. This, perhaps, is the key dimension worth examining more deeply and continually stress-testing going forward.

The chart below plots the software/services revenue mix against gross margin.

Source: Funda.ai

X-axis: software revenue mix; Y-axis: GAAP gross margin. The dashed line is the trend line. All figures are from each company’s most recent GAAP filings through March 2026.

In the top right are pure software companies — Adobe, Salesforce, Microsoft — with gross margins of 75–89%. Software has far lower marginal costs than hardware, and the leading software franchises enjoy pricing power through ecosystem lock-in; together, these factors push margins up. At the other end are hardware companies, specifically fabless semiconductor firms (no wafer fab, design only, manufacturing outsourced to TSMC): AMD 57%, Qualcomm 55%, Marvell 52%, MediaTek 46% — margins cluster in the 46–55% band. A fabless model entails a foundry-procurement cost for each chip sold, and this marginal cost structure determines the margin band.

In the middle sit hardware companies with a software ecosystem: Broadcom (software 35% of revenue, overall gross margin 68%), Cisco (software 51%, gross margin 65%), Arista (software 17%, gross margin 63%) — overall margins well above those of pure fabless peers. Strip software revenue out, however, and the hardware-only margins for all three fall back to 50–55%, on par with AMD and Qualcomm.

NVIDIA reports zero software revenue, yet carries a 75% gross margin and sits alone in the upper-left corner of the scatter plot. That premium comes from CUDA. CUDA is NVIDIA’s GPU programming platform, launched in 2001; code written in CUDA runs only on NVIDIA hardware. The tool itself is free, but over 25 years, it has attracted 6 million developers and 300+ acceleration libraries, with mainstream frameworks such as PyTorch relying heavily on closed-source libraries (cuDNN, cuBLAS) within the CUDA ecosystem for GPU acceleration, making migration extremely costly. Jensen Huang himself calls NVIDIA “a software platform company.” AMD — also fabless, selling to the same customer base, with its own software stack ROCm — has ecosystem maturity that falls well short of CUDA’s, and its gross margin is 57%. The 18-point gap is the CUDA moat.

But we are beginning to see marginal shifts worth taking seriously. DeepSeek’s upcoming V4 — continuing the company’s familiar playbook of “frontier performance + extreme cost efficiency” — will likely be one of the most closely watched open-source flagship models of 2026. According to public channels and press reporting, during the critical pre-release optimization window, DeepSeek allocated a substantial portion of that window preferentially to Huawei’s CANN ecosystem, while other platforms — including NVIDIA and AMD — did not receive comparable early access. The choice itself already reflects, at least in part, a trend toward ecosystem diversification.

What we think deserves more attention than any single vendor’s strategic choice is the change in the underlying “migration cost.” In the past, porting a frontier model from the CUDA stack to an alternative software/hardware stack (such as CANN) typically meant months of blocking engineering work for a team — a textbook high-friction process. With the introduction of AI coding agents, however, this kind of cross-stack adaptation is being materially accelerated, with cycle times potentially compressing from “quarters” to “weeks” or even “days.”

If this trend persists, the deepest assumption in the CUDA moat — that migration cost is high enough to produce structural lock-in — may be getting repriced. Put differently, the moat is not necessarily disappearing, but the mechanism that produced it is changing. Over the past two years, AI has been peeling back the moats of software companies layer by layer; the same mechanism begins to reach CUDA in 2026.


AI Is Dismantling Software’s Moats

IGV (the iShares Software ETF) is down 20% YTD through April 21, 2026, underperforming the S&P 500 by more than ten points.

The moats of software companies rest on lock-in: workflow lock-in, platform and data lock-in, and code-complexity lock-in. AI is peeling them back one by one.

Lightweight SaaS collapses first

AI hits “UI + simple logic” lightweight SaaS the hardest first — these products rely only on user habit and workflow inertia, so their lock-in is the shallowest. When AI agents replace the entire workflow, the lock-in simply disappears.

Monday.com (MNDY) is the canonical case: enterprise project management and task-collaboration SaaS, sold per seat. AI agents replace the seat directly, so enterprises no longer need to buy as many. MNDY started declining in mid-2025 and finished the year down 37%, while CRM and NOW were essentially flat. In 2026, the decline accelerated — management admitted the SMB self-serve channel had been persistently weak and would remain so for the full year, cut revenue growth guidance from 27% to 18–19%, and withdrew long-term targets. YTD -52% (through April 21).

Monday.com (MNDY) 1-year price chart, YTD -52%

Source: TradingView

Heavy enterprise software is no longer holding up

One layer deeper are heavy enterprise platforms — Salesforce and ServiceNow rely not just on workflow habits but also on years of accumulated customer data, custom configuration, and deep integration with organizational processes. These held up in 2025; in 2026, they began to take damage as well.

Salesforce (CRM): the world’s largest CRM platform; customer data, sales processes, and custom configurations from 150,000+ enterprise customers accumulated over the years all reside in Salesforce, making migration extraordinarily costly — the foundation of per-seat pricing. But AI agents can read customer data directly, execute follow-up and approval workflows automatically, and complete what used to require humans clicking through the Salesforce UI — bypassing the interface. Enterprises no longer need to buy a license for every sales rep and support agent. YTD -29% (through April 21).

Salesforce (CRM) 1-year price chart, YTD -29%

Source: TradingView

ServiceNow (NOW): enterprise IT workflow automation platform, with similar lock-in logic — ticket handling, change approvals, and IT operations all run on ServiceNow. AI agents can automatically classify tickets, approve requests, and trigger operational actions, reducing the platform’s reliance on human labor. Q4 earnings beat expectations, yet the stock dropped 10% that day; YTD -28% (through April 21).

ServiceNow (NOW) 1-year price chart, YTD -28%

Source: TradingView

Even “impossible-to-rewrite” COBOL has been cracked by AI

COBOL (Common Business-Oriented Language) was invented in 1959 — 67 years ago — specifically to process large-scale business data and transactions. The “foundation slab” of the global financial system is written in COBOL: 95% of ATM transactions and 80% of in-person credit card transactions run on COBOL systems; the core system at the U.S. Social Security Administration is 60M+ lines of COBOL code. An estimated 800B lines of COBOL code remain in production globally, processing more than $3 trillion in transactions every day.

This is precisely the core of IBM Z mainframe lock-in. BNY Mellon alone has 343M lines of COBOL code accumulated over 40–50 years. These programs handle money; changing one line and introducing a bug means billions of dollars of transactions gone wrong — no bank CEO wants to bear that risk. The average COBOL programmer is 55+ and retiring; young engineers do not learn COBOL, creating a hostage dynamic: the fewer people who can work on it, the less anyone dares touch it, and the more everyone depends on IBM for maintenance. These programs are deeply bound to IBM Z’s zOS operating system and cannot be simply lifted and shifted to the cloud.

Anthropic published a blog in February 2026 claiming that Claude Code can rapidly refactor COBOL code. IBM dropped 13% on the day — its largest single-session decline since October 2000 — and finished the month down 27%. The market had already rendered its verdict: once AI can write COBOL, IBM’s moat collapses instantly. CUDA’s moat is built on code complexity as well. Will it hold up?

How AI helps break the cost barrier to COBOL modernization — Anthropic Blog, 2026.2.23

Source: Anthropic

IBM 1-year price chart; -13% on the day of Anthropic’s COBOL blog in Feb 2026

Source: TradingView


Where Is CUDA’s Moat

CUDA is an entire compute stack, from GPU silicon up to the developer’s Python code, with an abstraction at every layer. When a developer calls torch.matmul (PyTorch’s matrix-multiply function), PyTorch calls a heavily optimized compute kernel inside cuBLAS (NVIDIA’s linear-algebra library); those kernels are written by NVIDIA engineers in CUDA C++, compiled via nvcc into GPU machine code, and dispatched at runtime by the driver to the GPU cores. From Python down to silicon, CUDA owns the entire path.

The CUDA Platform Stack Today — NVIDIA’s official CUDA full-stack architecture

Source: Nvidia

CUDA’s moat concentrates in three places: the outermost developer ecosystem, the Libraries & SDKs layer in the middle, written in CUDA C++, and the closed-source compiler below it.

Three layers of the CUDA ecosystem moat — ecosystem, closed-source libraries, closed-source compiler

Source: Nvidia

Ecosystem moat — the flywheel of 6 million developers

Jensen Huang himself states: “The single most important property of NVIDIA is the installed base of CUDA — the developer ecosystem.” Twenty-five years have yielded 6 million developers, 300+ acceleration libraries, teaching in 350+ universities, with GPU acceleration in mainstream frameworks like PyTorch depending primarily on the closed-source libraries cuDNN and cuBLAS inside the CUDA ecosystem. Huang’s flywheel narrative: developers write algorithms, algorithms spawn new applications, new applications expand the installed base, and the installed base attracts more developers. The flywheel has been spinning for 25 years. In his telling, 6 million developers are 6 million independent decision nodes, each depending on CUDA libraries, each accumulating CUDA code — migration cost spread across millions of nodes, no one able to coordinate an exit. That is the ecosystem network effect.

Closed-source library moat — closed-source kernels hand-written in CUDA C++

NVIDIA’s CUDA architect Stephen Jones gave an excellent breakdown at GTC 2025. His core point: parallel programming is “hard and annoying and difficult to debug”; CUDA built this entire layered platform precisely to keep developers from writing parallel code directly. Developers should stay at the highest abstraction they can — frameworks (PyTorch et al.) deliver results fastest; libraries (cuBLAS et al.) get close to peak performance; hand-writing kernels requires the greatest effort — and 99% of the parallelized operations that matter are already written in NVIDIA’s libraries, with only perhaps 1% requiring a developer to hand-write.

Pick Where To Spend Your Effort — Stephen Jones GTC 2025

Source: Nvidia

In other words, the most valuable part of the CUDA platform is the Libraries & SDKs layer in the middle. The libraries are the core acceleration kernels — cuBLAS (linear algebra), cuDNN (deep learning), cuFFT (Fourier transform), and others — essentially a set of compute kernels that NVIDIA engineers have spent nearly 20 years deeply optimizing in CUDA C++ for each successive GPU generation; the core libraries are all closed-source. SDKs package multiple libraries, tools, and APIs into domain-specific development kits — RAPIDS (data science), CUDA-Q (quantum computing), BioNeMo (drug discovery). Jones is extremely confident about this moat: the libraries are hand-tuned by NVIDIA’s top engineers (he calls them “ninjas”) over many years — “you would spend years to write the program that came close,” “these are the things that you should never even try and match,” “if you do write a program that beats these, you should definitely talk to me because I will hire you.” Even someone with 15 years of CUDA experience would take months to hand-write a kernel that matches cuDNN within 10%.

The logic of this lock-in is: the time of top-tier human engineers is a scarce resource. Every kernel in the library is the product of months or years of hand-tuning, and any competitor reproducing this library requires comparable magnitudes of human capital and time.

Closed-source compiler moat — ptxas, the required path for all code

Beneath the libraries lies another line of defense, hidden in the compiler.

Full view of the CUDA compilation pipeline (source: NVIDIA CUDA Compiler Driver NVCC Documentation 13.2, §3)

Source: Nvidia

This diagram shows the nvcc-driven offline compilation trajectory of CUDA: on the left is the host (CPU) path, on the right the device (GPU) path. The .cu source is split by nvcc into host/device parts; the host code is handed to a standard C++ compiler, while the device code first generates PTX, then is converted by ptxas into architecture-specific cubin, packaged as a fatbinary, and finally linked with host-side object files into the executable.

The PTX specification itself is public, and other toolchains do not necessarily need to pass through nvcc to enter the CUDA ecosystem. OpenAI’s Triton official documentation and NVIDIA’s own 2026 blog note that the Triton compiler can emit PTX; OpenXLA likewise states that XLA:GPU, on certain paths, invokes Triton as the PTX code-generation layer. But regardless of whether the front-end passes through nvcc, the back-end landing from PTX to target machine code remains controlled by NVIDIA.

In other words, a developer can bypass nvcc, but cannot bypass NVIDIA’s control over the PTX→target machine code stage. This “closed-source black box” is a critical segment of NVIDIA’s moat.

Yet in 2026, all three moats are loosening at the same time.


Where the CUDA Moat Faces Pressure

The timeline runs back to April 15, 2026. Dwarkesh Patel released a 103-minute deep-dive interview with Jensen Huang. The questioning was highly concentrated and direct, centered on four vectors: Anthropic’s shift to TPU, the contraction of the China market, hyperscaler progress on in-house silicon, and the trend among top-tier labs of bypassing cuBLAS via Triton.

User's avatar

Continue reading this post for free, courtesy of FUNDA.

Or purchase a paid subscription.
© 2026 FUNDA · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture