AI's Fact Recall Struggles Revealed: Key Findings from Google Research
In brief
- Google researchers have uncovered a critical issue in large language models (LLMs): they often get facts wrong not because they lack the information, but because they can't recall it.
- Using a new framework called knowledge profiling, the team analyzed how LLMs handle factual queries.
- They found that while models like Gemini3 and GPT-5 encode nearly all facts during training, many errors stem from their inability to retrieve these encoded details during inference.
- The study introduces WikiProfile, a benchmark with 2,150 Wikipedia facts paired with detailed questions to test encoding, recall, and recognition.
- Results show that most factual mistakes are "recall failures," akin to losing keys rather than having empty shelves.
- This distinction is crucial for improving LLM reliability-while encoding issues require bigger models or more data, recall problems may be solved with better retrieval methods during inference.
- Looking ahead, this research highlights the need for enhanced recall techniques in AI systems.
- Future advancements could focus on optimizing how LLMs access stored information, potentially leading to more accurate and trustworthy responses across various applications.
Terms in this brief
- knowledge profiling
- A method used by Google researchers to analyze how large language models handle factual queries, focusing on their ability to encode and retrieve information during training and inference.
- WikiProfile
- A benchmark created by Google researchers that consists of 2,150 Wikipedia facts paired with detailed questions. It tests the encoding, recall, and recognition abilities of large language models.
Read full story at Google AI Research →
More briefs
LLM Security Flaw Discovered Through Linguistic Illegibility
A new study reveals that large language models (LLMs) can produce outputs that don't accurately reflect their internal computations. This "linguistic illegibility" means relying on the model's own explanations or linguistic features to understand its operations is inherently unreliable. For instance, while an LLM might explain its reasoning in natural language, this explanation could fail to capture how it truly processes information, which is based on mathematical computations rather than language. This discovery has significant implications for security mechanisms designed to monitor and control AI systems. Current methods like chain-of-thought monitoring or self-critique may not be sufficient because they depend on the model's linguistic outputs. The researchers suggest using taint tracking-a method that identifies data influenced by the model-as a more reliable way to ensure safety. They also propose additional measures, such as robust virtualization and third-party audits, to create a stronger security framework. This research highlights the need for more advanced isolation techniques in AI systems to prevent potential exploits.
AI Revolutionizes Some Sciences, Falls Short in Others
A new study by Google, DeepMind, and MIT reveals that AI is making waves in fields like mathematics but struggling to make an impact in biology and drug discovery. While researchers across various disciplines are increasingly using AI tools, the technology's contributions remain limited. For instance, only 44% of scientists reported that AI has helped them save time, while many still face bottlenecks in experimentation and data collection. Despite government emphasis on AI's potential to transform science, practical applications like faster drug development have yet to materialize. The report highlights the gap between AI's hype and its real-world utility, urging further focus on areas where it can truly add value.
Six New AI Architectures Boost Semantic Search and Knowledge Graphs
Researchers have unveiled six cutting-edge architectures designed to enhance semantic search, knowledge graphs, and large language model (LLM) reasoning. These innovations aim to bridge the gap between understanding context and generating accurate answers. The systems are built for real-world applications, making them more reliable and efficient in processing complex queries. The new models integrate semantic search with knowledge graphs, allowing AI to better understand relationships between data points. This could lead to improved recommendation systems, smarter chatbots, and more effective information retrieval tools. By combining graph-based reasoning with LLMs, these architectures enable machines to answer questions based on both structured data and unstructured text. The research highlights the importance of practical applications over theoretical advancements. The architectures are already being tested in production environments, with early results showing significant improvements in query accuracy and response times. As AI continues to evolve, these patterns will likely influence future developments in natural language processing and knowledge management systems.
AI Gains Traction Among Industrial Chemists
At the Chemical Innovation Exchange (CIEX) conference in Indianapolis, artificial intelligence dominated discussions as companies showcased AI-driven tools for chemical discovery and lab operations. These tools, including those from Basetwo, Cypris, and Albert Invent, aim to reduce experimentation time, predict trends, and analyze data more efficiently than traditional methods. Chemists and executives are cautiously optimistic about AI's potential but remain skeptical of its accuracy compared to general-purpose models like GPT. The focus was on building trust in AI by highlighting its specialized applications. For example, Basetwo uses AI combined with physics-based insights to optimize chemical manufacturing processes, while Cypris leverages patent data to forecast technological trends. These tailored solutions offer concrete benefits, such as cost reductions and faster innovation cycles. However, the challenge lies in ensuring AI's reliability in chemistry-specific tasks, where even specialized models often fail. Moving forward, the integration of AI into chemical research is inevitable. Companies are investing in tools that complement human expertise, aiming to accelerate discovery while maintaining accuracy. As trust in these systems grows, AI is poised to become a critical partner for chemists in driving innovation across industries.
Google Hires Top Economists to Study AI's Impact on Jobs and Economy
Google has expanded its AI & Economy team by bringing in leading economists and researchers. This group will focus on understanding how artificial intelligence affects jobs, productivity, and global economic growth. The new directors, Anu Madgavkar and Daniel Rock, aim to analyze AI adoption worldwide, track labor market changes, and support business growth through their research. By using real-time data, they hope to assist both workers and policymakers in adapting to the evolving role of AI in the workforce. This initiative underscores Google's commitment to exploring how technology can benefit everyone, with findings available on ai.google/economy. Moving forward, this expanded team will continue to provide valuable insights to shape future policies and workforce development programs.