LLM Recall: When AI Knows the Fact but Still Misses the Answer

LLM Recall

Google published a report titled “Empty shelves or lost keys? Recall is the bottleneck for parametric factuality” by Nitay Calderon and Gal Yona, Research Scientists, Google Research, which provided a few groundbreaking loopholes in LLM functions. 

LLM Recall

This report says that sometimes LLMs give factually wrong answers even when they have learnt or stored and processed the information correctly.

Here is a concise account of the report with key conclusions. 

What the Report says about LLM Recall?

The report clearly signifies the difference between encoding failures and lost keys recall failures. 

It used an analogy for it, i.e., 

Empty Shelves = Encoding Failures

Lost Keys = Recall Failures.

So don’t get confused with the terms mentioned in the title of the report.  

They used knowledge profiling for this analysis, which is a behavioral framework that measures the 2 above mentioned aspects. 

In one line, the report says, 

“We then show that many factual errors in frontier LLMs are better understood as lost keys (recall failures), not empty shelves (encoding failures).”

And the whole analysis revolves around this inference. 

Key Conclusions 

The key conclusions extended by the report are supported by some stats. 

It said frontier models like Gemini-3-Pro and GPT-5 failed to recall 26–34% of facts, despite encoding 95–98% of facts. 

The report says:

“This means that in frontier models, factual errors increasingly come not from absent knowledge, but from knowledge that is stored and not reliably accessible.”

The study also analysed the results of Gemma 3 and concluded that its higher models show fewer encoding failures, but recall failures still have a significant share.

It also mentions that the rarity of the facts also has a certain impact on the encoding and recall functions of LLMs, and it comes with a correlation. 

They say that the difference between encodes of high- and low-popularity facts is very low, but for recalls the gap stretches significantly over tested LLMs, i.e., Gemini and GPT. 

According to the report, the models can encode around 90% of the long-tail facts, but the recall gap further slides by 20-35%

The same trend applies with the direct and reverse queries. The report presents an interesting case and calls it the Reversal Curse and explains it as: 

“When LLMs know “A is B” but can’t answer “What is B?”

They clearly infer that the LLMs explain in a specific order, and when the query order is reversed, it becomes difficult for these models to comprehend it. 

The report says, “The reversal curse is a recall problem,” which poses a concern and an opportunity that we will look into later in the discussion. 

LLMs use a bunch of strategies to derive the answer, such as Query Fanout, RAG, and chunking, and they keep on improving those approaches over time. So we can expect the reversal curse issue will also be rectified soon. 

Suggested Read: AI Citations vs. Brand Mentions: What AI Search Visibility Really Tells You

Opinions And Suggestions For Digital Marketers

Chris Westmeyer, CEO and Founder of Digital Strike, has provided certain valuable suggestions in his LinkedIn post that SEOs can adapt to yield the opportunity stated in the report. He said

  • Write the reverse sentence somewhere on the page. Lead with the category, not the brand name.
  • Use FAQ blocks to state key facts in both directions.
  • Audit your About and Service pages. Most of them run brands first all the time.

Also, Dan Petrovic, a renowned Australian AI and SEO expert, linked it to AI hallucination that happened even when the encoding goes right.

LLM Recall opinions

Final Words

From this LLM recall research by Google and the recent Reddit’s ChatGPT citation drop, it is evident that the nature of LLMs is unpredictable, and approaching it requires smart SEO strategies. 

However, considering these revelations, it is advised to make minor tweaks in your approach, rather than completely overhauling the entire work plan. 

Next Read: Why Does AI Search Prefer Original Research Citations? Possible Gaps and Solutions

Previous Article

How AI Agents for SEO Can Turn Data Into Action

Next Article

JSON-LD Escaping: What Google’s Parser Change Means for SEO

Write a Comment

Leave a Comment

Your email address will not be published. Required fields are marked *

Subscribe to our Newsletter

Subscribe to our email newsletter to get the latest posts delivered right to your email.
Pure inspiration, zero spam ✨