Google Bets on Specialized Chips to Boost AI Model Efficiency
تعمل جوجل على تطوير شريحة خوادم جديدة تحت الاسم الرمزي Frozen v2، تهدف إلى تحسين كفاءة معالجة استفسارات نموذج جيميناي وخفض استهلاك الطاقة، وسط تزايد الطلب على خدمات الذكاء الاصطناعي.
Google is developing a new server chip dedicated to running its AI model 'Gemini', in a step aimed at increasing request processing efficiency and reducing energy consumption, amid growing pressure on the company's computing infrastructure and accelerating demand for AI services.
According to two reports published by Bloomberg and The Information, the new chip carries the codename Frozen v2, and is expected to adopt a different approach by integrating parts of Gemini models' underlying architecture into the hardware itself, rather than relying entirely on software to perform all processing operations. This design aims to speed up response and improve resource consumption efficiency, while reducing the load on servers.
In a comment to CNBC, Google said its engineering teams are continuously researching and testing new technologies, with the goal of achieving the highest levels of performance and efficiency for users and customers.
The company added that testing any new project does not necessarily mean it will reach commercial production stage, indicating that the Frozen v2 project may remain within the scope of research if it does not achieve the targeted results.
For its part, a Google Cloud spokesperson explained that the company adopts an approach based on designing hardware and software in an integrated manner from the early stages of development, allowing for improved system performance and alignment with actual workloads, especially large-scale AI applications.
Integrating Gemini into the Chip: The Frozen v2 concept relies on integrating some components of Gemini's architecture directly into the silicon chip, reducing the number of computational operations the chip must perform each time a user sends a request to the model.
In traditional chips, data moves repeatedly between memory and processing units, while multiple stages of computational operations are performed before generating the final answer. The more distance data travels and the more steps required to process it, the higher the energy consumption and response time.
In the new chip, Google seeks to fix some parts of the model's architecture within the physical components themselves, allowing for faster execution of specific operations, reducing data movement within the system, and lowering the computational overhead associated with generating responses.
This means, for example, that instead of requiring the chip to reload certain instructions or model components for each request, these components can be pre-existing within the chip's design and ready for direct execution.
Higher Efficiency: Engineers working on the project estimate that the Frozen v2 chip could achieve an increase of between 6 and 10 times in the number of text units (tokens) that can be produced per watt of power, compared to the latest generation of Google's custom AI processing units known as TPUs.
A 'token' refers to a part of a word, a whole word, or a punctuation mark that the AI model processes during response generation. The number of tokens the system can produce within a specific time period is one of the key indicators for measuring model speed and operational efficiency.
Thus, increasing the number of text tokens produced per watt means Google can process more requests using the same amount of electricity, or provide the same performance with lower energy consumption.
This efficiency is becoming increasingly important as Gemini expands into search, email, cloud services, and enterprise products, as operating AI models at scale requires massive amounts of electricity and data centers equipped with thousands of chips.
The reports indicated that the new chip will not replace the general-purpose TPUs currently used by Google, but will represent a specialized branch within the company's chip portfolio.
TPUs can run various types of AI models and workloads, while Frozen v2 will be more closely designed around Gemini's own architecture.
This degree of specialization gives the new chip higher efficiency, but in return imposes important technical constraints: if Google decides to significantly change the underlying architecture of Gemini models, the components embedded in the chip may become incompatible with newer versions.
Consequently, Frozen v2 will not be able to support future versions of Gemini unless Google maintains the same core components in the model's architecture.
Crisis in Computing Capacity: Google developed the project partly to address an growing internal crisis in computing capacity, which has led to friction between different company divisions, as teams compete for the resources needed to train and operate AI models.
According to reports, those pressures have driven Google Cloud to reject some jobs and requests from external customers, due to insufficient computing power to meet growing demand.
Major technology companies face a dual challenge: providing enough chips to train new models, while simultaneously running existing models for hundreds of millions of users and businesses.
Operating models after training, known as 'inference', is one of the most resource-intensive processes in the long run, because each request sent by a user to Gemini requires a series of computational operations to generate a response.
By developing a chip specialized for Gemini's architecture, Google hopes to reduce the cost of this stage and increase the number of requests that can be processed within data centers.
2028 Launch Target: Google aims to start using the new chip during 2028, but the project is still in development, and engineers have not yet finalized the design.
The company is still studying how much of Gemini's model information can be permanently embedded in the chip, versus parts that should remain flexible and modifiable via software.
This point is considered one of the main challenges facing the project, because integrating large parts of the model into silicon gives the chip higher efficiency, but reduces its ability to adapt to future changes in Gemini's architecture.
It is expected that production volumes of Frozen v2 will be much lower than those of TPUs, as Google currently treats the project as an exploratory experiment, not a program ready for broad commercial launch.
This means the company may first use the chip for specific tasks within its data centers to gauge its ability to reduce energy consumption and accelerate Gemini operations, before deciding on production expansion.
Success of the project could also lead to the development of other generations of chips dedicated to specific models, rather than relying on general-purpose chips that can be used with a wide range of systems.
Race for Specialized Chips: Google is not the only company moving toward developing custom AI chips; major companies are racing to reduce their dependence on external chip manufacturers and lower the operational costs of their models.
OpenAI recently introduced its first internally developed AI chip in collaboration with Broadcom, in a step aimed at providing a more specialized alternative for running its models and services.
Original source: Asharq News
Comments (0)
Be the first to comment.