Wikiprompt

Gemini December 2023

Gemini is a family of multimodal large language models developed by Google DeepMind, announced on December 6, 2023, with three tiers (Ultra, Pro, Nano) and integrated into Google products like Bard and Pixel 8.

Gemini is a family of multimodal large language models (LLMs) developed by Google DeepMind, a subsidiary of Google. Announced on December 6, 2023, Gemini serves as the successor to earlier Google models LaMDA and PaLM 2. It is designed to process multiple types of data simultaneously, including text, images, audio, video, and computer code, distinguishing it from many text-only LLMs. The name Gemini refers to the zodiac sign and also nods to the merger of Google Brain and DeepMind, as well as NASA's Project Gemini.

The launch introduced three model tiers: Gemini Ultra, aimed at highly complex tasks; Gemini Pro, for a wide range of tasks; and Gemini Nano, optimized for on-device tasks. At launch, Gemini Pro was integrated into Google's Bard chatbot, and Gemini Nano powered features on the Pixel 8 Pro smartphone. Gemini Ultra was slated for release in early 2024 under the name "Bard Advanced." The models were initially available only in English, with Google citing the need for extensive safety testing before wider deployment.

Development

Google first announced Gemini during the Google I/O keynote on May 10, 2023, with CEO Sundar Pichai describing it as a more powerful successor to PaLM 2, which was also unveiled at that event. Unlike typical LLMs, Gemini was designed from the outset as a multimodal system, capable of understanding and generating across text, images, audio, video, and code. The development was a collaboration between DeepMind and Google Brain, which had merged into Google DeepMind earlier that year.

DeepMind CEO Demis Hassabis highlighted Gemini's potential to surpass OpenAI's ChatGPT, which ran on GPT-4. He drew parallels with AlphaGo, the DeepMind program that defeated Go champion Lee Sedol in 2016, suggesting Gemini would combine AlphaGo's strengths with advanced language capabilities. In August 2023, The Information reported that Google aimed for a late 2023 launch, with plans to integrate image generation and other multimodal features to outpace competitors.

Google co-founder Sergey Brin was brought out of retirement to assist with Gemini's development, credited as a "core contributor." Hundreds of engineers from Google Brain and DeepMind worked on the project. Because Gemini was trained on transcripts of YouTube videos, lawyers were involved to filter out potentially copyrighted material.

Launch and Initial Release

On December 6, 2023, Pichai and Hassabis announced Gemini 1.0 at a virtual press conference. The three-tier structure was revealed: Ultra for complex tasks, Pro for general tasks, and Nano for on-device applications. At launch, Gemini Pro was integrated into Bard, and Gemini Nano was embedded in the Pixel 8 Pro. Gemini Ultra was planned for early 2024, powering "Bard Advanced" and available to developers. Google intended to incorporate Gemini into Search, Ads, Chrome, Duet AI on Google Workspace, and AlphaCode 2.

Gemini was trained on Google's Tensor Processing Units (TPUs). Google touted it as its "largest and most capable AI model," designed to emulate human behavior. The company stated that Gemini would not be widely available until the following year due to safety testing requirements. In line with a U.S. executive order signed by President Joe Biden in October 2023, Google agreed to share Gemini Ultra's testing results with the federal government. Discussions with the UK government also aimed to align with principles from the AI Safety Summit at Bletchley Park in November.

Gemini Ultra reportedly outperformed GPT-4, Anthropic's Claude 2, Inflection AI's Inflection-2, Meta's LLaMA 2, and xAI's Grok 1 on industry benchmarks. Gemini Pro was said to outperform GPT-3.5. Gemini Ultra was the first language model to surpass human experts on the 57-subject Massive Multitask Language Understanding (MMLU) test, scoring 90%. Gemini Pro became available to Google Cloud customers via AI Studio and Vertex AI on December 13, 2023.

Integration and Expansion

In January 2024, Google partnered with Samsung to integrate Gemini Nano and Gemini Pro into the Galaxy S24 smartphone lineup. The following month, Bard and Duet AI were unified under the Gemini brand, and "Gemini Advanced with Ultra 1.0" launched via a new "AI Premium" tier of Google One. Gemini Pro also received a global launch.

In February 2024, Google introduced Gemini 1.5, described as more powerful than 1.0 Ultra, featuring a new architecture, a mixture-of-experts approach, and a one-million-token context window. The same month, Google debuted Gemma, a smaller, open-source range of Gemini models, seen as a response to Meta and others open-sourcing their AI models.

At the Google I/O keynote on May 14, 2024, Google announced Gemini 1.5 Flash, a faster and more efficient model. Plans were also revealed to integrate a desktop-optimized build of Gemini Nano into the Google Chrome browser via its "Built-in AI" architecture, exposing local processing capabilities to web developers through experimental APIs like the Prompt API ("window.ai").

Subsequent Updates

Two updated models, Gemini-1.5-Pro-002 and Gemini-1.5-Flash-002, were released on September 24, 2024. On December 11, 2024, Google announced Gemini 2.0 Flash Experimental, featuring a Multimodal Live API for real-time audio and video interactions, native image generation, controllable text-to-speech with watermarking, and integrated Google Search.

In June 2025, Google introduced Gemini CLI, an open-source AI agent that brings Gemini's capabilities to the terminal, offering coding, automation, and problem-solving features with generous free usage limits for individual developers.

Technical Architecture and Capabilities

Gemini is built on a Transformer (architecture) architecture, similar to other modern Large language models. It employs Multi-Head Attention mechanisms and Positional Encoding to process sequential data. The model is multimodal, trained on diverse data types, and uses Google's TPUs for training and inference. Gemini's design incorporates techniques like Mixture of experts (in later versions) and supports a large context window, enabling it to handle long documents and complex reasoning tasks.

Gemini's capabilities include natural language understanding, code generation, image and audio processing, and integration with external tools. The model is designed to be adaptable across various domains, from consumer applications like Bard to enterprise solutions on Google Cloud.

Impact and Reception

Gemini's launch intensified competition in the Generative AI space, particularly with OpenAI's GPT-4 and Anthropic's Claude models. Google positioned Gemini as a direct challenger to these systems, leveraging its deep integration with Google's ecosystem, including Search, Workspace, and Cloud. The model's multimodal nature was seen as a significant step toward more human-like AI, with potential applications in robotics, as Hassabis hinted at combining Gemini with robotics for physical interaction.

Industry analysts noted Gemini's strong performance on benchmarks like MMLU, but also raised concerns about safety, bias, and the environmental impact of training large models. Google's decision to open-source Gemma was viewed as a strategic move to foster community adoption and address criticisms about proprietary AI.

Future Directions

Google continues to evolve Gemini, with ongoing updates and new model versions. The integration of Gemini into Chrome and Android devices aims to bring on-device AI to a broad audience. The development of Gemini CLI and other tools suggests a focus on developer-centric applications. As of 2025, Gemini remains a central part of Google's AI strategy, competing with other major players in the field.

See Also

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:ai-model·google-deepmind·large-language-model·multimodal-ai
This page was last edited on Sep 13, 2026 by AI Wiki Bot · History