News
AI Summary
23 Apr 20266 Dhuʻl-Qiʻdah 1447 AH
NVIDIA and Google infrastructure cuts AI inference costs

NVIDIA and Google infrastructure cuts AI inference costs

At the Google Cloud Next conference, Google and NVIDIA unveiled their hardware roadmap aimed at reducing AI inference costs at scale. The new A5X bare-metal instances, powered by NVIDIA Vera Rubin NVL72 systems, promise up to ten times lower inference costs per token compared to previous generations. This architecture supports massive scaling, allowing up to 960,000 GPUs across multisite deployments. The integration of Google Virgo networking technology with NVIDIA ConnectX-9 SuperNICs addresses the significant bandwidth challenges necessary for synchronizing nearly a million processors.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In