Google’s “Frozen v2” Chip Could Make Gemini 10x More Efficient

frozen v2 chip

Short Answer

Google is developing a new server chip internally called Frozen v2 that hardwires parts of Gemini’s model architecture directly into silicon. Engineers project it could process 6 to 10 times more AI tokens per unit of power than Google’s current TPUs. The chip is targeted for deployment as early as 2028 and sent Alphabet stock up roughly 3% on the news.

Quick Summary

  • Google is building a chip called Frozen v2 that embeds Gemini’s model architecture directly into the hardware, reducing the computation needed to generate AI responses
  • Engineers project 6 to 10 times more tokens per unit of power compared to Google’s current-generation TPUs
  • Unlike TPUs, which work with many models, Frozen v2 is built exclusively for Gemini, trading flexibility for efficiency
  • Alphabet stock rose roughly 3% on the report, which was first published by The Information on July 20, 2026
  • Deployment is targeted for as early as 2028, and Google says Frozen v2 will complement rather than replace its existing TPU lineup

Alphabet shares jumped roughly 3% on Monday after CNBC reported that Google is developing a new AI chip called Frozen v2 designed to run its Gemini models at dramatically lower energy costs, citing a report by The Information. The report landed at a moment when the cost of serving AI at scale has become one of the central tensions in the industry: the more capable the model, the more expensive every response.

Frozen v2 takes a different approach to that problem. Rather than building a faster general-purpose chip, Google’s engineers are proposing to lock Gemini’s model structure into the hardware itself. The name follows the logic of “freezing” parameters in AI training, where you stop certain values from changing. Applied to silicon, it means some of Gemini’s decision-making is physically built into the chip rather than computed on the fly.

For brands and content teams tracking Answer Engine Optimization, the implication runs downstream: cheaper inference means more AI answers served, more queries routed through Gemini, and more opportunities for content to be cited or surfaced. A chip that makes Gemini 10x more power-efficient is also a chip that makes Gemini cheaper to run at scale.

The NumbersWhat It Means
6–10xMore tokens per unit of power compared to Google’s current TPUs, according to Google engineers’ projections
3%Alphabet stock gain on July 20, 2026 following The Information’s report on the Frozen v2 chip
2028Earliest projected deployment date for Frozen v2 in Google’s server infrastructure
1 modelFrozen v2 is built exclusively for Gemini, unlike general-purpose TPUs that run across many models

What Frozen v2 Actually Is

In one line: Frozen v2 is a purpose-built chip that hardwires Gemini’s architecture into silicon to cut the energy cost of running AI responses.

Most AI chips, including Google’s own TPUs, are general-purpose accelerators. They are designed to handle any model you throw at them, which requires flexibility. Flexibility costs power. Frozen v2 trades that flexibility for efficiency by baking Gemini’s specific architecture into the chip’s physical design. Parts of the model that would normally be computed at runtime are instead fixed in the hardware.

The name is deliberate. In AI training, “freezing” a layer means stopping it from updating further. You lock in the values because they are good enough and you do not want them to change. Frozen v2 applies that idea at the hardware level: Gemini’s structure is locked into the silicon. The result is a chip that cannot run other models well but can run Gemini faster and with less power than anything currently available.

Google describes the broader strategy as co-designing hardware and software “from the ground up,” a phrase it has used for its TPU program for years. Frozen v2 takes that philosophy further than TPUs do, tying a specific model to a specific chip rather than optimizing a chip family for a class of workloads.

Why Efficiency Matters More Than Raw Speed

In one line: The AI industry’s near-term constraint is not model capability but the cost of serving responses at scale, and Frozen v2 directly targets that constraint.

The race in AI hardware has shifted. A year ago the question was which chip could train the biggest model fastest. Today the more pressing question is which chip can serve the most responses for the least electricity, because that is what determines the economics of running AI products at consumer scale. Power costs are the single largest variable in data center operating expenses, and they are rising as demand for AI inference grows.

A chip that delivers 6 to 10 times more tokens per unit of power does not just save money. It changes the math on what is profitable to offer for free, what can be run on cheaper hardware, and how many parallel requests a data center can handle. Google has repeatedly said that making AI responses cheaper is central to its cloud and consumer strategy. Frozen v2, if the projections hold, is the most direct bet it has placed on that outcome.

What this means for AEO: When serving an AI response costs less, AI products get more aggressive about answering questions rather than linking out. More answers served through Gemini means more chances for your content to be cited, and more chances for brands that have not structured their content for AI citation to get left out. The efficiency curve rewards preparedness. See where to start: how to improve your Answer Engine Optimization.

What Comes Next

In one line: Frozen v2 is a 2028 target, not a product announcement, but its existence signals that Google is planning for Gemini to be its primary AI surface for years.

Google has not officially announced Frozen v2. The report came from The Information, citing unnamed people familiar with the project, and Google’s response to CNBC was careful: the company confirmed it is “constantly researching and experimenting with new innovations” without confirming the chip by name. That is standard for hardware at this stage, where timelines and specs shift frequently before tape-out.

The 2028 deployment target puts Frozen v2 roughly two chip generations out from today. Between now and then, Google will continue shipping TPU generations and updating Gemini. Frozen v2 would slot in alongside those products, not replace them. TPUs stay for multi-model flexibility; Frozen v2 handles the Gemini inference load that now dominates Google’s AI serving.

For investors, the 3% stock move reflected something beyond the chip itself. It signaled that Google has a credible hardware path to making its AI business more cost-efficient, which matters at a moment when the company is spending heavily on infrastructure and being asked by analysts to show a return. Bloomberg also covered the report, noting the market reaction as a signal of investor confidence in Google’s hardware roadmap. A chip that cuts the per-token cost of Gemini by an order of magnitude is a significant answer to that question.

Key Takeaways

  • Gemini baked into silicon: Frozen v2 hardwires Gemini’s architecture into the chip itself rather than computing it at runtime, trading model flexibility for dramatic efficiency gains
  • 6 to 10x more efficient: Google engineers project that figure in tokens per unit of power versus current TPUs, though these are internal projections, not published benchmarks
  • Not a TPU replacement: Frozen v2 would run alongside Google’s general-purpose TPU lineup, handling Gemini inference specifically
  • 2028 at the earliest: The chip is mid-roadmap, not imminent. Google has not confirmed it publicly by name
  • Cheaper inference = more AI answers: For brands, a more efficient Gemini means more queries answered by AI, which raises the stakes for content structured to be cited in those answers

Frequently Asked Questions

What is Google’s Frozen v2 chip?

Frozen v2 is a server chip Google is reportedly developing that embeds parts of its Gemini AI model’s architecture directly into the silicon. By hardwiring Gemini’s structure into the hardware, the chip reduces the computation needed to generate responses and could deliver 6 to 10 times more tokens per unit of power than Google’s current TPUs.

How does Frozen v2 differ from Google’s TPUs?

Google’s TPUs are general-purpose AI accelerators that can run many different models. Frozen v2 is designed exclusively for Gemini. By giving up flexibility, it gains significant efficiency: parts of Gemini’s model that TPUs compute at runtime are instead built into the chip’s physical design.

When will Frozen v2 be available?

According to The Information’s July 2026 report, Google is targeting deployment as early as 2028. The chip has not been officially announced and timelines for hardware at this stage typically shift before production.

Why did Alphabet stock rise on this news?

Alphabet shares gained roughly 3% because the report suggested Google has a hardware path to making Gemini significantly cheaper to run. Reducing per-token power costs at the scale Google operates would meaningfully improve the economics of its AI products and cloud business, which investors have been watching closely given the company’s heavy infrastructure spending.

What does this mean for AI assistants and search?

More efficient inference means Google can serve more AI responses for the same energy cost. That makes it economically viable to answer more queries through Gemini rather than returning traditional links, which increases the importance of structuring content to be cited in AI answers rather than just ranked in search results.

About the author

Kai Williams

Kai Williams has been in marketing for years, with a long background in SEO before AEO had a name. He stepped into Answer Engine Optimization the moment AI started reshaping how people search, and has been tracking the shift ever since. At Prompt Insider, he covers AEO, AI marketing, and the future of search, breaking down what is changing and what brands need to do about it.

Keep reading

Be a Prompt Insider. Get AI news, AEO insights, resources, and updates delivered straight to your inbox.