Google is developing a new server chip specifically designed to run its intelligent model, Gemini, in a move aimed at improving request processing efficiency and reducing energy consumption, amid growing pressure on the company's computing infrastructure and accelerating demand for AI services.

According to reports by Bloomberg and The Information, the new chip is codenamed Frozen v2 and is expected to adopt a different approach by integrating parts of Gemini's architecture into the hardware itself, rather than relying entirely on software to perform all processing tasks. This design aims to speed up response times, improve resource efficiency, and reduce server load.

In a comment to CNBC, Google said its engineering teams continuously research and test new technologies to achieve the highest levels of performance and efficiency for users and customers.

The company added that testing any new project does not necessarily mean it will reach commercial production, indicating that the Frozen v2 project may remain within the research scope if it does not achieve the desired results.

A Google Cloud spokesperson explained that the company adopts an approach of integrated hardware and software design from the early stages of development, allowing system performance optimization and alignment with actual workloads, especially large-scale AI applications.

Embedding Gemini inside the chip

The idea behind Frozen v2 is to integrate some components of Gemini's architecture directly into the silicon chip, reducing the number of computational operations the chip must perform each time a user sends a request to the model.

In traditional chips, data frequently moves between memory and processing units, while multiple stages of computations are performed before producing the final answer. The greater the distance data travels and the more processing steps required, the higher the energy consumption and response time.

In the new chip, Google aims to embed certain parts of the model's architecture into the physical components themselves, allowing specific operations to be executed faster, reducing data movement within the system, and lowering the computational overhead associated with generating responses.

This means, for example, that instead of requiring the chip to reload certain instructions or model components for each request, these components can already be present in the chip's design and ready for immediate execution.

Higher efficiency

Engineers working on the project estimate that the Frozen v2 chip could achieve a 6- to 10-fold increase in the number of tokens produced per watt of energy, compared to the latest generation of Google's custom AI processing units, known as TPUs.

A token refers to a part of a word, a whole word, or a punctuation mark that the AI model processes when generating responses. The number of tokens the system can produce in a given time is a key metric for measuring model speed and efficiency.

Thus, increasing the number of tokens produced per watt means Google can handle more requests using the same amount of electricity, or deliver the same performance with lower energy consumption.

This efficiency becomes increasingly important as Gemini expands into search, email, cloud services, and enterprise products, since running AI models at scale requires massive amounts of electricity and data centers equipped with thousands of chips.

The reports indicated that the new chip will not replace the general-purpose TPUs Google currently uses, but will represent a specialized branch within the company's chip portfolio.

TPU chips can run various types of AI models and workloads, while Frozen v2 will be designed more closely tied to Gemini's own architecture.

This degree of specialization gives the new chip higher efficiency, but it also imposes significant technical constraints: if Google decides to significantly change Gemini's underlying architecture, the components embedded in the chip may become incompatible with new generations.

Therefore, Frozen v2 will only be able to support future versions of Gemini if Google maintains the same core components in the model's architecture.

Crisis in computing capacity

Google developed the project partly to address a growing internal crisis in computing capacity, which has led to friction between different divisions of the company, as teams compete for resources needed to train and run AI models.

According to reports, these pressures have forced Google Cloud to reject some work and requests from external customers due to a lack of sufficient computing power to meet growing demand.

Major technology companies face a dual challenge: providing enough chips to train new models while simultaneously running existing models for hundreds of millions of users and businesses.

Running models after training, known as inference, is one of the most resource-intensive processes in the long term, because each request a user sends to Gemini requires a series of computations to generate an answer.

By developing a chip specialized for Gemini's architecture, Google hopes to reduce the cost of this stage and increase the number of requests that can be processed in data centers.

2028 target for launch

Google aims to begin using the new chip in 2028, but the project is still in development, and engineers have not yet finalized the design.

The company is still evaluating how much of Gemini's model information can be permanently integrated into the chip, versus parts that should remain flexible and adjustable via software.

This is one of the key challenges facing the project, because integrating large parts of the model into silicon gives the chip higher efficiency, but reduces its ability to adapt to future changes in Gemini's architecture.

Production volumes of Frozen v2 are expected to be much smaller than those of TPUs, as Google currently treats the project as an exploratory experiment, rather than a program ready for large-scale commercial launch.

This means the company may first use the chip for specific tasks within its data centers to measure its ability to reduce energy consumption and speed up Gemini performance, before deciding on expanding production.

The project's success could also lead to the development of other generations of chips dedicated to specific models, rather than relying on general-purpose chips that can be used with a wide range of systems.

Race for specialized chips

Google is not the only company moving toward developing custom AI chips; major companies are racing to reduce their dependence on external chip manufacturers and lower the costs of running their models.

OpenAI recently introduced its first in-house AI chip developed in collaboration with Broadcom, a move aimed at providing a more specialized alternative for running its models and services.