Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)

/ 5:46 / E610

What Does “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)” Talk About?

This episode of the Fatrank Podcast dives into one of the most important yet misunderstood concepts in AI visibility: the difference between parametric and non-parametric memory in large language models. Host James Dooley interviews Stephen Burns from Common Crawl to unpack how LLMs like ChatGPT, Anthropic, and Gemini decide whether to recall an answer from their trained weights or perform a live RAG (retrieval augmented generation) search. Stephen explains that parametric memory is the content baked into the model's weights before its cutoff date, delivered fluently and confidently without citations, while live search kicks in when the model doesn't recognise a brand, product, or when it knows fresher information exists.

“This is the content that was in the training data before the model's cutoff date. It's baked into the weights.”

— Stephen Burns

Who Are the Guests on “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)”?

James Dooley is the host of the Fatrank Podcast and a well-known figure in the SEO and AI visibility space. He brings a practitioner's perspective, drawing on his experience working with agencies and his focus on emerging strategies like getting brands baked into LLM training data rather than only chasing traditional Google rankings.

What Are the Key Takeaways From “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)”?

Here are the key points discussed in this episode:

  • LLMs use two types of memory: parametric memory baked into the model's weights, and live RAG search for content the model doesn't recognise or needs to update.
  • To get into parametric memory, SEOs should build high harmonic centrality links from sites close to the core of the web, such as Wikipedia and major news sites.
  • Getting content into the parametric memory of an LLM can take six months to over a year due to crawl download and machine learning cycles.
  • SEOs now work in two channels: optimising for parametric memory through authority links, and optimising on-page content for quick RAG searches.
  • When analytics show fast results within a month, that indicates a RAG search rather than parametric memory, which helps interpret AI visibility performance.

“Well, the data shows that it can take six months to over a year to get into the parametric memory.”

— Stephen Burns

Is “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)” Worth Listening To?

For anyone doing SEO or AI visibility work in 2026, understanding the distinction between what LLMs remember and what they look up is a genuine competitive edge. The discussion around harmonic centrality, Common Crawl's role in the training corpus, and the practical advice to work in two channels simultaneously makes this a short but dense listen packed with strategy you can apply immediately.

Who Should Listen to “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)”?

This episode is ideal for:

  • SEO professionals adapting their strategy for AI and LLM visibility
  • Digital marketing agency owners exploring AI search optimisation
  • Brand managers wanting to appear in ChatGPT, Gemini, and Anthropic answers
  • Technical marketers curious about Common Crawl and harmonic centrality

Where Can You Listen to Fatrank Podcast?

You can listen to Fatrank Podcast on all major podcast platforms:

  • Apple Podcasts – Search for “Fatrank Podcast” in the Podcasts app
  • Spotify – Available on Spotify for free
  • Amazon Music / Audible – Listen through your Amazon account
  • Overcast – For iOS users who prefer a dedicated podcast app
  • Pocket Casts – Cross-platform podcast player

You can also subscribe using the RSS feed: https://feeds.transistor.fm/fatrank-podcast

What Are Listeners Saying About This Episode?

★★★★★

“Finally someone explained the difference between parametric memory and RAG search in plain English. The point about results showing up within a month meaning it was a live search totally changed how I read my analytics.”

— Daniel R-K.

★★★★★

“Stephen Burns from Common Crawl brought real depth here. The advice on building harmonic centrality links from Wikipedia and news sites to get baked into training data is exactly the kind of practical insight I needed.”

— Priya S-M.

★★★★★

“Short episode but incredibly dense. Learning that it can take six months to a year to get into parametric memory reset my expectations and helped me think about SEO as two separate channels now.”

— Marcus T-L.

This video explains which digital marketing strategies SEO and AI visibility agencies should focus on in 2026 to improve LLM visibility, brand trust and long term search presence. James Dooley and Stephen Burns start with KPI tracking because analytics reveal whether results come from parametric memory or a live RAG search, which changes how you interpret performance. They cover brand SEO, AI visibility and Google Business Profiles because stronger search presence improves trust and conversion rates.
The discussion also explores organic SEO, organic social media and paid social ads because consistent visibility across search and social supports long term growth. PPC is analysed in detail because campaign setup, landing pages and lead handling directly affect results. They also discuss Reddit, Quora and paid AI ads because diversified enquiry sources and early adoption can strengthen digital marketing performance for SEO and AI visibility agencies.
PromoSEO lead generation for SEO and AI visibility agencies recently received recognition as the “Best SEO And AI Visibility Agencies Lead Generation Agency.”
Where to Listen to This Episode
Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data is available on:

James Dooley: The two memories of AI, you've got live search which is performed by RAG and you've got parametric memory, which Common Crawl is actually part of the training corpus. And today I'm joined with Stephen Burns from Common Crawl. So to kick things off, can we talk about parametric memory and what that actually means?

Stephen Burns: Sure. That's the first layer, parametric memory. This is the content that was in the training data before the model's cutoff date. It's baked into the weights. And when the model answers a question that touches this content, the answer comes back fluent. It comes back fast and confident. There's no citations. The model just recalls it from memory.

James Dooley: And then with regards to that, so I'm just going to read a couple of things off here. So what LLMs remember and what they look up obviously is the two kind of difference. So at what point is it where, okay, I've now got this part of the training corpus versus, oh, I need to go and look that up, perform retrieval augment if generate thingy, perform RAG basically, to go and perform the live search? Why sometimes they need to go and do a live search?

Stephen Burns: It's going to do a live search when it looks... It's going to look in its memory. Do I know what this is? Do I know what this product is? Do I know this brand? And if it doesn't know what it is, then it's going to do the live search. Or if it... It may make a decision. You can make a decision, say, well, there's new... I know there's new information on this. Let's also get the current information on this product or brand.

James Dooley: And then obviously as SEOs of the world and people are looking to try to increase AI visibility or LLM visibility, that could be in ChatGPT or Anthropic or Gemini. Ideally now what you should be doing is not just trying to look to rank better within Google, but actually try and start to get baked in to that training data. So, can you explain how Common Crawl, if you increase the CC Rank via harmonic centrality, how that can actually help you get part of the training corpus?

Stephen Burns: So, yeah, you're going to want to get good, uh, high harmonic centrality links to your site. Those sites that are close to the core of the web, those are usually the most popular brands. Uh, you're looking at Wikipedia links, you're looking at news sites, high-end, you know, big news site links, those types of sites. That's going to get you seen more often into the parametric memory.

James Dooley: And then with regards to the parametric memory, because there's quite a lot of people that don't understand properly on there, would you say that it's important for SEOs to be looking at doing both? Because I see a lot of people talking about consensus and trying to get rankings in Bing and trying to get rankings in Google to try to get into the AI Overviews or AI Mode or ChatGPT. How important and how long does it take if you're trying to get part of the training data for the LLMs to try to start picking up and updating the training corpus?

Stephen Burns: Well, the data shows that it can take six months to over a year to get into the parametric memory. You know, the crawl comes out and then the LLM may download it a month, a couple months later. And then when they do their next learning, it can take month, six months for them to actually do all the machine learning to learn it all and then finally publish it. So, it's, it's behind at least a year. So you have to also think as an SEO and go how, you know, you're, you're working in two channels now. You're saying you're going to be working trying to get into that memory. That's some of your work working on, uh, using HC and, and, and getting that ranking. And then cit... Uh, you know, your, your search retrieval, your quick searches that are done inside the LLM. You're going to, you know, work on content on your pages for that or, or other, other ways of doing that.

James Dooley: Yeah, you're going to notice when you start doing analytics that, you know, some people say, "Well, we changed the page and made it more crawlable and we updated this content. How come it's not showing up?" Well, sometimes it may not show up because it hasn't gotten into the parametric memory. And number two, um, and if it does show up right away and you notice a, a result, you know, within a, within a month, you're like, "Wow." Well, that's because probably because it was a, a, a search, a RAG search.

Stephen Burns: Yeah, for sure.

James Dooley: Anyone who's watching this and you're now starting to understand and maybe dig a bit deeper into parametric memory for LLMs, make sure you check out the link in the description. I do several different episodes with Stephen Burns from Common Crawl. I personally met him out in Vietnam in Saigon and I was amazed because I didn't actually realise that Common Crawl was so important for LLM visibility. Another thing is one of the episodes talks about AI visibility audits. Make sure you check that out to see whether you're not blocking CCBot or any of the LLM bots that are out there. There's one also where we can kind of dig deep on the algorithms behind Common Crawl, which uses harmonic centrality. Stephen Burns, it's been an absolute pleasure. Thanks for having you and I appreciate everything here of you talking because I didn't properly understand the terminology of parametric memory. I knew I heard part of the training data or performing a live search, but the two different memories and how they started to do it, it's...

Stephen Burns: Yeah, it's, in... It's, for me, it's intriguing. Um, I'm always looking to try to increase AI as much as I can with the visibility. So, thanks for having you.

Creators & Guests

James Dooley Host
James Dooley

James Dooley is the founder of FatRank which is a UK lead generation company. James Dooley is the current CEO of FatRank that provides high-quality leads for UK business owners.

Stephen Burns Guest
Stephen Burns

Stephen Burns is a technical SEO and generative engine optimisation specialist with 25 years of experience in search. As Web Intelligence Lead at the Common Crawl Foundation and Principal Technical…

No episode selected
0:00
0:00