# Business Insights

By [DYLIT Chronicles](https://dylit.info/user/dylitmediabuzz)

[Everything AI](https://dylit.info/pr/everything-ai/6a9efac02e92664f4d50cf9d) > [Business Insights](https://dylit.info/ch/business-insights/6a9efac02e92664f4d50cfac)

D-Matrix Adopts NVIDIA's Chip-Linking Technology for AI Servers What D-Matrix Just Agreed To Santa Clara startup d-Matrix announced this week that it will build support for NVIDIA's NVLink Fusion interconnect into its upcoming Raptor chips, letting them connect directly to NVIDIA's data center server racks. The arrangement folds NVLink Fusion, an interconnect NVIDIA has licensed to outside chipmakers since last year, into d-Matrix's future chip designs, including the upcoming Raptor accelerators. Neither company disclosed financial terms. D-Matrix says the combined systems target fast, low-latency AI services such as coding assistants, chatbots and voice agents, where speed is critical. d-Matrix builds chips for inference, the work of running an already-trained model to answer a live user request, a different job from the training workloads NVIDIA's GPUs have traditionally dominated. What NVLink Fusion Actually Does NVLink Fusion is NVIDIA's system for letting chips made by other companies talk to NVIDIA GPUs and servers at very high speed. NVIDIA describes it as high-bandwidth, low-latency technology and IP that lets hyperscalers and AI companies plug custom processors and CPUs into NVIDIA's broader AI infrastructure platform, built around the MGX rack-scale architecture. Before this, NVLink's fast interconnect was reserved for NVIDIA's own silicon. NVIDIA opened it up in 2025, releasing hardware and IP designed to let third-party CPUs and accelerators interoperate with NVIDIA chips over NVLink connections. In d-Matrix's case, the technology gives Raptor chips specialized connectors and memory so they can sit inside the same rack as NVIDIA hardware, rather than running as a separate, incompatible system. Why This Matters for Both Companies For d-Matrix, the deal solves a real engineering problem. Building NVLink Fusion into its designs helps the startup sidestep many of the difficulties involved in scaling its own chip architecture across large compute clusters on its own. Founded in 2019, d-Matrix has raised around $500 million so far and reached roughly a $2 billion valuation, with Microsoft investing through its M12 venture arm. For NVIDIA, the arrangement works differently. Even if a customer picks d-Matrix silicon over an NVIDIA GPU for a given inference job, that workload still runs inside NVIDIA's rack architecture and interconnect fabric. NVIDIA gains either way, whether the chip doing the work carries its logo or not. The Competitive Ripple Effect D-Matrix isn't the first or only company to sign onto NVLink Fusion. MediaTek, Marvell, Alchip Technologies, Astera Labs, Synopsys and Cadence were among the earliest adopters, and Fujitsu and Qualcomm are integrating their own custom CPUs with NVIDIA GPUs through the same program. Amazon's chip unit is also building on the platform for upcoming Trainium4 deployment. That growing list matters competitively: it suggests NVIDIA is positioning itself as the connective layer underneath the AI chip market, rather than competing chip-for-chip against every rival. Rival inference chip makers, including Groq, Cerebras and SambaNova, are chasing the same fast-inference market without this kind of formal tie to NVIDIA's rack ecosystem, a difference that could shape how each company's hardware gets adopted into existing data centers. What This Means for Customers, and What's Still Unclear For AI companies running large-scale inference, the appeal is compatibility: a d-Matrix-NVIDIA rack could, in theory, let operators mix specialized inference silicon into infrastructure already built around NVIDIA hardware, without standing up an entirely separate system. But important details remain unverified. Neither company has published independent benchmarks comparing Raptor's performance or cost efficiency against NVIDIA's own inference GPUs. At the Hot Chips conference, d-Matrix showed early results running two AI models on a rack of Raptor chips, reaching about 3,000 tokens per second per user on one model and 1,000 on another, with the company saying this scales across multiple racks. Those figures come from d-Matrix itself, not an independent test. The combined systems aren't expected until 2027, leaving time for the technology and the competitive landscape to shift before any of this reaches production.   Sources NVIDIA Newsroom — https://nvidianews.nvidia.com/news/nvidia-nvlink-fusion-semi-custom-ai-infrastructure-partner-ecosystem The Register — https://www.theregister.com/systems/2026/09/10/d-matrix-drinks-the-nvidia-kool-aid-with-nvlink-fusion-and-mgx-rack-designs/5295403 The Next Platform — https://www.nextplatform.com/compute/2026/09/10/startup-d-matrix-will-pair-its-raptor-memory-based-xpu-to-nvidia-rackscale-iron/5295601 YouTube Videos "NVIDIA NVLink Fusion: Building Semi-Custom AI Infrastructure" — https://www.youtube.com/watch?v=BLLf4e9BqXs "Understanding NVIDIA NVLink" — https://www.youtube.com/watch?v=6lm8zXvfjyc
