Google Launches New AI Model for Multimodal Computing
In brief
- Google has introduced Gemma 4 12B, a groundbreaking multimodal AI model designed to run efficiently on laptops.
- This model eliminates the need for separate encoders for vision and audio, instead processing both inputs directly within its architecture.
- This innovation reduces latency and memory usage while maintaining high performance comparable to larger models.
- The new model is particularly notable for being the first mid-sized model with native audio support, making it versatile for tasks like speech recognition and image analysis.
- It requires only 16GB of VRAM, enabling seamless operation on standard laptops.
- Developers have already used earlier Gemma models to create applications ranging from robotic arms to AI security systems.
- This release marks a significant step in bringing advanced AI capabilities to everyday devices without compromising speed or functionality.
- As Google continues to refine its multimodal approach, we can expect even more powerful and accessible tools for developers and users alike.
Terms in this brief
- Gemma 4 12B
- A multimodal AI model developed by Google that processes both visual and audio inputs directly within its architecture. It's efficient enough to run on laptops with only 16GB of VRAM, making it accessible for various applications like speech recognition and image analysis.
Read full story at DeepMind Safety →
More briefs
SK Hynix Sees 1242 Percent Net Profit Boost
SK hynix said its second-quarter net profit soared 1242 percent year-on-year. This was driven by demand for its memory chips from the artificial intelligence industry. The company's quarterly net profit was 94 trillion won, an all-time high. Operating profit jumped 557 percent from last year to 60 trillion won. Revenue stood at 79 trillion won. SK hynix will invest around 40 trillion won this year. The company expects demand for its memory chips to persist as tech companies increase their AI infrastructure investments. SK hynix will continue to grow with the evolving AI technology.
UK Introduces AI-Enabled Smart Lamp-Posts
The UK has introduced AI-enabled smart lamp-posts that can recognize number plates and faces. These lamp-posts can power themselves through solar panels and have the potential to fight crime and track down missing persons. They can also incorporate new AI capability such as gait analysis and suspicious behaviors. The use of these lamp-posts has raised concerns about surveillance and data ownership, with 50,000 of them set to be installed in Nigeria. The technology will continue to develop and expand in the future.
AI Companies Destroying Millions of Books for Training Data
AI companies are destroying millions of physical books to use for training data. They scan the books and then discard them. This matters because the books are a valuable source of human-authored text. The AI companies need this text to improve their models. One company, Anthropic, was sued for copyright infringement and paid a $1.5 billion settlement. The AI companies are now hiring middlemen to buy the books for them. They want to keep their involvement a secret because they know it is not popular. The demand for human-authored text will continue to drive the destruction of physical books.
AI Sees Through Leaves to Help Farmers
Scientists created a new AI technology that can see through leaves to identify and measure hidden fruit. This technology can help farmers by creating a complete 3D model of each plant, including what is behind the leaves. It can open the door to automating tasks like crop forecasting. Current methods can produce inaccuracies of up to 23%, but the new technology has achieved fruit counts within 2% to 3% of the correct figure. The technology can save large growers millions of dollars and reduce waste. It will help farmers understand how much fruit they will produce and the size of the fruit. Farmers will be able to plan better with this new technology.
Microsoft's New Cybersecurity Model Delivers Big Efficiency Gains
Microsoft has unveiled MAI-Cyber-1-Flash, a compact cybersecurity model that achieves an impressive 96 percent score on the CyberGym benchmark when integrated into its MDASH multi-agent system. This new tool significantly reduces costs by cutting down reliance on expensive pure frontier models-costs are expected to drop by 50 percent since only the most challenging cases will be handled by GPT-5.4. While MAI-Cyber-1-Flash excels in efficiency, Microsoft still turns to OpenAI for complex reasoning tasks. This hybrid approach allows Microsoft to leverage its own model where it shines brightest while relying on OpenAI's expertise for tougher problems. Looking ahead, this cost-effective solution could expand cybersecurity capabilities for businesses, making advanced protection more accessible. The integration with MDASH suggests a broader push toward multi-agent systems in security, potentially leading to even smarter and more coordinated defenses in the future.