During its annual Cloud Next conference held this week in Las Vegas, Google Cloud announced the launch of the Virgo network, an advanced networking architecture specifically designed to connect hundreds of thousands of AI accelerators within an integrated and efficient computing scope.
The company described this pioneering system as an approach that transforms the entire technology campus into a single massive computer, replacing traditional public data center networks with a custom-designed mesh fabric tailored to the immense and complex demands of training and operating advanced AI models. This approach reflects a deep realization that generative AI and large language models require unprecedented infrastructure to operate at full efficiency.
A three-dimensional network architecture enhancing flexibility
The new architecture separates data center networks into 3 distinct and independent layers, enabling tremendous flexibility in managing information flow. The first layer is dedicated to elevating connection levels within a single pod, while the second layer acts as a fabric dedicated to transferring data from one accelerator to another across different pods. The third layer is a frontend interface connecting to Google’s existing Jupiter network, ensuring seamless access to storage resources and general computing resources.
This ingenious engineering design allows each layer to be upgraded independently, sparing the entire system from any disruptions or downtime when updating or developing a component, thereby ensuring continuity for business and research projects without costly pauses.
Reducing latency and efficiently isolating faults
At its core, this advanced system relies on high-capacity switches linked to a two-tier flat design that does not obstruct data flow, which Google emphasizes contributes significantly to reducing latency by lowering the number of network tiers compared to traditional legacy designs.
Furthermore, the multi-tier design featuring independent control domains provides the advantage of effectively isolating faults. In the event of any hardware malfunction within a single tier, the issue is immediately contained so that the entire system does not collapse and the failure does not spread to the rest of the interconnected network.
Massive capabilities in connecting processors
Using its new TPU v8t chips, the advanced Virgo network can connect approximately 134,000 processors with a massive bandwidth of up to 47 petabits per second inside a single data center fabric. When scaling up to include multiple data centers across different geographic locations, this staggering figure rises to over one million processors all connected within a massive cloud training network.
Regarding setups based on Nvidia GPU processors, the system supports up to 80,000 graphics processors in a single location, reaching 960,000 processors when linking multiple locations together. These figures reflect the positive transformative power of this system in accelerating training processes for giant AI models.
Precise monitoring and doubled performance
Statistics revealed by the company indicate that the new system delivers 4 times the bandwidth per accelerator compared to the previous generation, alongside a 40 percent reduction in latency. Innovations do not stop there; the new architecture also includes precise telemetry technologies operating in fractions of a millisecond, alongside automated detection of lagging or unresponsive network nodes. This feature is critical for protecting AI training tasks from localized failures that could waste thousands of computing hours.
Comprehensive overhaul of cloud infrastructure
This announcement came as part of a broader technology infrastructure renewal plan showcased at Cloud Next 2026, which ran from April 22 to 24 at the Mandalay Bay complex. Alongside the network, the company launched its eighth-generation AI processors, divided into training-focused TPUs (TPU 8T) and inference- and operation-focused TPUs (TPU 8I).
- Introduction of revolutionary innovations in Managed Lustre cloud storage.
- Provision of data storage throughput capabilities reaching up to 10 terabits per second.
- Unlimited support for handling expanding AI model parameters.
In a technical post detailing the new system, the company stated clearly, quoting:
“The era of AI requires a radical rethinking of physical cloud infrastructure, and networking in particular.”
The company emphasized that as the size of AI models grows exponentially, traditional networks face inevitable breakdown points regarding bandwidth, responsiveness, and the ability to handle the massive and simultaneous data flows that characterize large-scale AI training operations today, making this innovation an urgent necessity for the future.
Frequently Asked Questions
What is the prominent technology recently announced by Google Cloud?
The company announced a new and innovative networking system aimed at connecting hundreds of thousands of AI accelerators in an integrated computing environment, functioning as a connected supercomputer.
How does the new network system handle technical faults and issues?
The system relies on a multi-tier design with independent control domains, allowing faults to be isolated with absolute precision within a single tier and preventing them from spreading to affect other parts of the network or data center.
How many processors can the system connect simultaneously?
The system can connect about 134,000 processors in a single data center, and this number exceeds one million processors when connecting multiple data centers across different geographic locations to build a giant training network.
Where and when were these technological innovations announced?
These innovations were unveiled during the annual Cloud Next 2026 conference, held in Las Vegas from April 22 to 24.