What is Google TPU v7 Ironwood and why does it change AI?

Last update: November 27th 2025
  • Google TPU v7 Ironwood is the seventh generation of TPUs and is specifically designed for the inference era, with 4.614 TFLOPs FP8 and 192 GB of HBM3e per chip.
  • The architecture scales up to 9.216 chips in a superpod with 1,77 PB of shared memory, 9,6 Tb/s ICI networks, and liquid cooling optimized for nearly 10 MW.
  • Integrated into AI Hypercomputer along with Axion CPUs, Ironwood offers a strong balance between performance, energy efficiency and cost for massive AI workloads.
  • The mega-contract with Anthropic and the vertical integration strategy reinforce Google's role as a key competitor to Nvidia in the AI ​​chip market.

Google TPU v7 Ironwood

The acceleration of artificial intelligence is undergoing a transformation . It's no longer enough to train massive models and boast about their parameters: now the real battle is how to make them work daily, for millions of users, without AI energy and infrastructure costs skyrocketing. In this context, Google TPU v7 Ironwood emerges, a piece of hardware designed for an AI that not only responds, but also thinks, reasons, and acts continuously.

Ironwood is Google's seventh-generation Tensor Processing Unit (TPU) and, according to the company itself, its most powerful, scalable, and efficient chip to date. It's designed for the so-called "era of inference," where what matters isn't so much how long it takes to train a new model, but how much it costs to serve it, keep it active, reasoning, and generating results around the clock. Let's take a closer look at exactly what Google TPU v7 Ironwood is, what makes it special, and why it's shaking things up against Nvidia and the other giants in the industry.

What is Google TPU v7 Ironwood and why is it so important?

Ironwood is essentially a custom-built AI accelerator developed by Google to run cutting-edge artificial intelligence models. It belongs to the family of Tensor Processing Units (TPUs), chips specifically designed for machine learning workloads that have been powering internal Google services (search, YouTube, advertising, translation, etc.) for nearly a decade, and more recently, Google Cloud solutions.

This seventh generation has been built upon all the experience accumulated with previous TPUs and with a clear objective: to maximize performance in inference and complex reasoning tasks , without sacrificing trainability. Large Language Models (LLMs), Mixture of Experts, augmented recall systems, and AI agents that need to maintain a lot of context in memory are the types of workloads Ironwood is designed to handle.

Instead of simply being “a faster chip,” Ironwood is integrated into Google Cloud’s AI Hypercomputer architecture , an approach where computing, networking, storage, and software are designed together. This allows for fine-tuning from the silicon to the container orchestrator to squeeze every watt and every clock cycle.

Another key point is that Ironwood is the first TPU designed from the ground up for the inference era . Until now, most AI systems were primarily designed to train massive models; now the focus is shifting to running those models efficiently, repeatedly, with agents that not only react but also anticipate, search for data, combine it, and return conclusions.

TPU v7 Ironwood Architecture

Key technical features of TPU v7 Ironwood

On the technical front, Ironwood boasts impressive figures, even in a market as competitive as AI chips. Each chip delivers 4.614 teraflops of FP8 precision computing power , which is more than 16 times the computing capacity of the TPU v4 and a huge leap forward compared to previous generations.

To power this immense processing capacity, each chip integrates 192 GB of next-generation HBM3e memory , with an approximate bandwidth of 7,3-7,4 TB/s. This combination of high-speed computing and memory is precisely what makes it possible to handle gigantic models and very long contexts without experiencing bottlenecks.

Ironwood is manufactured using an advanced 5-nanometer process and consumes approximately 600 watts per chip. This may sound high, but when compared to its performance per watt, the result is very competitive: compared to the previous generation Trillium (TPU v6e), Ironwood doubles the performance per watt.

These improvements are noticeable in both training and service workloads. For training, Ironwood can accelerate the processing of next-generation models (such as Gemini or Veo) and those from other vendors, and for serving, it enables massive inferences to be run with lower energy costs, something critical now that AI data centers are measured in gigawatts.

Superpod Ironwood

From chip to Superpod: how Ironwood scales

The true magic of Ironwood lies not in a single chip, but in how it scales to form AI supercomputers . Each Ironwood SoC is mounted on a PCBA with four chips, and each rack integrates 16 of these boards, for a total of 64 chips per rack.

  How to use DeepMind and understand its real impact on AI

From there, the racks are grouped together to form an Ironwood Superpod . A complete superpod includes 144 racks, which equates to 9.216 interconnected chips. In total, this translates to around 42,5 exaflops of FP8 performance and approximately 1,77 petabytes of shared HBM memory, all accessible as a vast logical resource for AI models.

The interconnection between chips is achieved through the Inter-Chip Interconnect (ICI) network , with a 3D Torus (4x4x4) design and speeds of up to 9,6 Tb/s. This topology allows for reduced latency and maintains high transfer rates even when the system scales to thousands of accelerators, which is essential for distributed training and highly parallel inference workloads.

Beyond the superpod, the infrastructure can be further extended by combining blocks of 64 connected chips and linking them via an optical backbone of up to 1,8 PB/s. This mix of passive copper cables, fiber optics, and optical switches provides enormous flexibility when configuring custom clusters for different customers and models.

In terms of power, a deployment of this size approaches 10 megawatts of electrical capacity , a figure that necessitates taking cooling and efficiency very seriously. Each rack is designed to exceed 100 kW of consumption, so the entire system has been designed from the ground up for liquid cooling and advanced power distribution.

Ironwood TPU Cooling Infrastructure

Refrigeration, consumption and energy efficiency

With such monstrous power densities, the obvious question is: how does Google keep all this under thermal control? The answer lies in a state-of-the-art liquid cooling system, designed specifically for the Ironwood superpods.

Each rack integrates a coolant distribution unit (CBU) with leak detectors and 416V AC electrical redundancy . This configuration not only prevents overheating but also provides high reliability for critical AI workloads, where a prolonged outage could result in millions of euros in losses or large-scale service disruptions.

Ironwood isn't just aiming to be the most powerful system, but also one of the most efficient. Compared to Trillium, the TPU v7 achieves twice the performance per watt consumed . This means that, with the same 2024 energy budget, a data center equipped with Ironwood could handle up to twice the inference by 2026.

This approach addresses an uncomfortable reality in the sector: energy, not silicon, is becoming the limiting factor for AI expansion. The electrical grids of many countries are beginning to reach their capacity in key data center areas, so extracting more useful work from each watt is no longer an advantage, it's a necessity.

Furthermore, Ironwood's architecture is optimized to minimize unnecessary data movement, a major energy consumer in distributed systems. The combination of massive HBM memory and high-speed networks aims precisely at that: more processing closer to the data and fewer costly trips across the data center.

AI Hypercomputer Google Cloud

Ironwood within AI Hypercomputer and the software layer

Ironwood doesn't exist in isolation; it's deeply integrated into Google Cloud's AI Hypercomputer architecture . This concept brings together hardware (Axion CPUs and TPUs), high-speed networking, storage, and orchestration and scheduling software within a single design framework.

At the software layer, Google relies on its Pathways stack, which allows for the coordination of tens of thousands of Ironwood TPUs as if they were a single logical resource. Alongside Pathways, tools like MaxText (for large-scale LLM training ) and integration with Kubernetes Engine via Cluster Director have been strengthened , optimizing cluster scheduling and resilience.

For the inference component, Google has extended vLLM support to work with both GPUs and TPUs , making it easier to migrate models and workloads between different types of accelerators with minimal code changes. Furthermore, GKE Inference Gateway can reduce initial request latency by up to 96% and lower the cost of service by around 30%, according to data shared by the company.

All of this means that customers not only get a fast chip, but a complete, end-to-end optimized platform . IDC, in fact, cites returns on investment of 353% over three years for AI Hypercomputer customers, with IT cost reductions of nearly 28%.

The ultimate goal is clear: to enable businesses to harness the power of Ironwood without the hassle of managing complex infrastructures. The ideal user simply defines their models and needs, and the Google Cloud control plane seamlessly orchestrates TPUs, Axion, networking, storage, and software.

  Why Windows doesn't recognize the USB drive and how to fix it

TPU Ironwood in data centers

The “age of inference” and the role of Ironwood

For years, much of the AI ​​discourse revolved around training increasingly larger models . However, the market has been confronted with a harsh reality: no matter how spectacular the training, sustained revenue comes from providing inferences at scale, day in and day out.

Google refers to this new era as the “era of inference .” Instead of merely reactive models that wait for the user to ask questions, we are moving toward AI agents that retrieve data, analyze contexts, coordinate subprocesses, and take the initiative to offer recommendations and processed information.

Ironwood is designed precisely for these types of workloads: models with complex reasoning, extensive contexts, and persistent states . The large amount of shared memory and network bandwidth are intended to allow multiple components of the same AI system to work in a coordinated manner without becoming overwhelmed by traffic.

In this scenario, Google adopts a different strategy than Nvidia. While Blackwell chips continue to dominate cutting-edge training and promise very high inference speeds in specific configurations, Google focuses on optimizing the continuous and massive execution of models in production , integrating hardware and cloud into a single, coherent package.

The company doesn't intend to replace Nvidia across the board, but rather to capture a growing share of the overall inference market . Some internal estimates suggest that, by 2027, Google's TPUs could represent up to 30% of that market, especially among large clients already using Google Cloud.

The agreement with Anthropic and the commercial boost

Ironwood's major commercial boost came from Anthropic, the developer of the Claude models . On October 23, 2025, the company signed a massive agreement with Google, committing to use up to one million TPUs and invest tens of billions of dollars in infrastructure on Google Cloud.

This contract anticipates energy consumption exceeding 1 gigawatt by 2026 , an enormous figure that gives an idea of ​​the scale at which the major players in AI are operating. Overnight, Ironwood went from being an ambitious project on the roadmap to a real investment backed by strong demand.

For Google, the agreement resolves one of its biggest concerns: building data center capacity without guaranteed usage . With Anthropic as its anchor client, the company can plan new facilities, power contracts, and Ironwood superpod deployments with much greater visibility into their return on investment.

Anthropic's choice of Ironwood, rather than waiting for Nvidia's next generation or resorting to proprietary hardware from other hyperscalers, is a clear sign of confidence in the efficiency and time-to-market of Google's TPUs . "More inference, less power" is, in practice, the implicit motto of this move.

These types of long-term agreements also help Google to better negotiate its own costs (from chip manufacturing to energy supply) and strengthen its position in a market where the demand for AI computing far exceeds the available supply.

Impact on the AI ​​chip market and competition with Nvidia

The launch of Ironwood is a game-changer in the AI ​​hardware market , which has been clearly dominated by Nvidia GPUs until now. Google isn't the only one making moves: AWS is pushing its Trainium2 line, and Microsoft is progressing with its Maia program, albeit with delays.

Nvidia continues to lead in cutting-edge training, with its Hopper, Blackwell, and the upcoming Rubin architectures, which target configurations of several exaflops per rack. However, Google is pushing the market toward a model of multiple vertically integrated ecosystems , where each hyperscaler designs its own chips, network, software, and cloud services.

In this new landscape, Ironwood positions itself as a solid alternative for companies seeking large-scale inference infrastructure without relying entirely on the CUDA ecosystem. The combination of custom TPUs and Axion CPUs, plus the Google Cloud platform, creates a highly attractive offering for companies that require predictable costs, performance, and availability.

The move also adds pressure on Nvidia, which must justify high prices in an environment where it is no longer the only viable option . As Amazon, Microsoft, and Google bolster their own stacks, the market is shifting away from the single-chip vendor model toward a patchwork of closed, but highly optimized, solutions within each cloud provider.

However, the success of Ironwood and TPUs in general will depend heavily on Google's software stack continuing to mature at the same pace as the CUDA ecosystem . If cluster utilization rates drop (for example, from 90% to 70%), overall efficiency will suffer and price competitiveness could be affected.

  Google launches Gemma 3: its new open AI optimized for a single GPU

Axion: the CPU add-on for AI infrastructure

Alongside the Ironwood launch, Google has strengthened its Axion family of custom CPUs , based on the Arm Neoverse architecture. While TPUs handle the heaviest aspects of machine learning, Axion takes care of the general-purpose workloads surrounding AI.

The new N4A instances offer up to 64 vCPUs, 512 GB of DDR5 memory, and 50 Gbps network bandwidth , with up to 100% price-performance improvements compared to equivalent x86 machines. Meanwhile, C4A Metal is Google Cloud's first bare-metal version based on Arm, geared towards scenarios such as Android development, automotive applications, and advanced simulations.

Companies like Vimeo have reported a 30% improvement in video transcoding , while ZoomInfo reports a 60% increase in efficiency in data analysis processes. Other companies, such as Rise, claim to have reduced computing power consumption by around 20% while maintaining stable latency.

The underlying idea is for Axion to act as a general-purpose computing base layer : data preparation, microservices, web services, business logic, etc. On that foundation, Ironwood TPUs focus on what they do best: accelerating the training, inference, and reasoning of AI models.

This combination reinforces the vision that next-generation computing will be hybrid : efficient CPUs to manage workflows and TPUs (or other accelerators) for intensive computation. All of this is brought together by the AI ​​Hypercomputer platform and Google Cloud tools.

Investment outlook and margin protection for Alphabet

From a financial perspective, Ironwood is much more than a powerful chip: it's a strategic component for protecting Alphabet's margins in the cloud AI business. Hyperscalers are on track to reach double-digit gigawatt power capacities by 2027, and all indications are that infrastructure investment will remain massive.

By designing its own chips, Google can capture value across the entire value chain : from silicon design to cloud deployment. Instead of reselling third-party hardware with slim margins, it transforms capital expenditure (CAPEX) into a competitive advantage and a tool for differentiating its services.

Ironwood's improved performance per watt means that each megawatt of data center capacity delivers significantly more in terms of inferences served . Combined with software optimizations (vLLM, Pathways, GKE Inference Gateway, etc.), this allows Google to offer aggressive pricing without compromising profitability.

The contract with Anthropic adds an extra layer of stability to Google's capital spending plan . Instead of building infrastructure "blindly," waiting for customers to arrive, it can now size much of its expansion based on long-term agreements, reducing uncertainty and the risk of overcapacity.

Even so, open questions remain: whether Google will attract more large anchor clients, whether the energy projects and new substations will stay on schedule, and whether its software stack will maintain high utilization rates. The potential is there, but execution will be key for Ironwood to be both a technical success and an engine of sustained profits . THIS IS NOT INVESTMENT ADVICE.

Taken together, Google TPU v7 Ironwood represents a generational leap that goes far beyond the teraflops on a spec sheet: it combines raw power, massive shared memory, advanced optical networks, liquid cooling, and a deeply integrated cloud to support the new wave of AI agents and models running around the clock. In a world where energy is tight, models are constantly growing, and the competition for control of the AI ​​stack is fierce, Ironwood is emerging as the heart of an infrastructure designed to usher in the "age of inference" and redefine how, where, and at what cost artificial intelligence is run on a global scale.

What is tensorflow-0
Related articles:
What is TensorFlow and how it revolutionizes artificial intelligence