Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)
Listen on your favourite platform
| Platform | Link |
|---|---|
| YouTube | Listen on YouTube → |
| Transistor | Listen on Transistor → |
| pod.link | Listen on pod.link → |
| Pocket Casts | Listen on Pocket Casts → |
| Podverse | Listen on Podverse → |
| Castro | Listen on Castro → |
| Listen Notes | Listen on Listen Notes → |
| Rephonic | Listen on Rephonic → |
| getpodcast | Listen on getpodcast → |
| Spotify | Listen on Spotify → |
| Amazon Music | Listen on Amazon Music → |
| Castbox | Listen on Castbox → |
| Podcast Addict | Listen on Podcast Addict → |
| Podchaser | Listen on Podchaser → |
| Steno.fm | Listen on Steno.fm → |
What Does “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)” Talk About?
This episode of the Fatrank Podcast dives into one of the most important yet misunderstood concepts in AI visibility: the difference between parametric and non-parametric memory in large language models. Host James Dooley interviews Stephen Burns from Common Crawl to break down how LLMs like ChatGPT, Anthropic, and Gemini decide whether to recall information baked into their training weights or perform a live RAG (retrieval augmented generation) search for current data.
The conversation explains that parametric memory is content that existed in the training data before a model's cutoff date, delivered fast, fluent, and without citations. In contrast, a live search happens when the model doesn't recognize a brand or product, or when it knows newer information exists. Stephen and James discuss why Common Crawl is such a critical piece of the training corpus and how improving your CC Rank through harmonic centrality links from sites like Wikipedia and major news outlets can help you get 'baked into' that parametric memory.
They also cover the practical timelines and analytics implications for SEOs. Getting into parametric memory can take six months to over a year, while a quick ranking improvement is often the result of a RAG search rather than actual training data inclusion. The episode encourages SEOs to work in two channels simultaneously, optimizing both for long-term training data presence and for immediate retrieval-based visibility.
“This is the content that was in the training data before the model's cutoff date. It's baked into the weights. And when the model answers a question that touches this content, the answer comes back fluent. It comes back fast and confident. There's no citations. The model just recalls it from memory.”
— Stephen Burns
Who Are the Guests on “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)”?
James Dooley is the host of the Fatrank Podcast and a well-known figure in the SEO and digital marketing industry. He explores emerging strategies in AI visibility, LLM optimization, and search presence, bringing curiosity and practical questioning to complex technical topics on behalf of the SEO community.
Stephen Burns represents Common Crawl, the massive open web crawl dataset that forms a significant part of the training corpus used by large language models. His expertise lies in how content gets included in LLM training data, the role of harmonic centrality and CC Rank in web crawling, and how brands and SEOs can increase their AI visibility. James met Stephen in Saigon, Vietnam, and credits him with clarifying his understanding of the two distinct memory systems that power modern AI.
What Are the Key Takeaways From “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)”?
Here are the key points discussed in this episode:
- AI models operate with two types of memory: parametric memory baked into the model's weights from training data, and non-parametric live search performed through RAG.
- An LLM performs a live search when it doesn't recognize a brand or product, or when it knows newer information exists that it should retrieve.
- Getting into parametric memory can take six months to over a year because of the lag between crawls, downloads, and machine learning cycles.
- High harmonic centrality links from sites close to the core of the web, like Wikipedia and major news sites, help you get into parametric memory.
- SEOs should work in two channels at once, optimizing for both long-term training data inclusion and immediate RAG-based retrieval visibility.
“Well, the data shows that it can take six months to over a year to get into the parametric memory.”
— Stephen Burns
Is “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)” Worth Listening To?
This episode is valuable because it demystifies a concept that many SEOs hear about but few truly understand. The clear distinction between parametric memory and live RAG search reframes how you interpret analytics results, explaining why some content updates show up instantly while others take a year or more to appear in LLM responses. This insight alone can save agencies from misreading their performance data.
Beyond the theory, the conversation delivers actionable guidance. Stephen Burns explains the practical steps of building high harmonic centrality links, understanding Common Crawl's role in the training corpus, and running AI visibility audits to ensure you aren't blocking CCBot or other LLM crawlers. For anyone trying to future-proof their search strategy for 2026 and beyond, this episode offers a rare, source-level look at how getting baked into LLM training data actually works.
Who Should Listen to “Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data (James Dooley interviews Stephen Burns)”?
This episode is ideal for:
- SEO professionals looking to expand into AI and LLM visibility
- Digital marketing agency owners planning strategies for 2026
- Brand marketers wanting to appear in ChatGPT, Gemini, and Anthropic answers
- Technical marketers curious about how Common Crawl and training data work
Where Can You Listen to Fatrank Podcast?
You can listen to Fatrank Podcast on all major podcast platforms:
- Apple Podcasts – Search for “Fatrank Podcast” in the Podcasts app
- Spotify – Available on Spotify for free
- Amazon Music / Audible – Listen through your Amazon account
- Overcast – For iOS users who prefer a dedicated podcast app
- Pocket Casts – Cross-platform podcast player
You can also subscribe using the RSS feed: https://feeds.transistor.fm/fatrank-podcast
What Are Listeners Saying About This Episode?
“The breakdown of parametric memory versus RAG search finally made sense of why some of my content updates never seem to show up in ChatGPT. Stephen's point about the six month to over a year delay explained so much about my analytics. Genuinely eye-opening episode.”
“I had no idea Common Crawl was so central to LLM visibility until I heard this. The advice on getting high harmonic centrality links from Wikipedia and news sites is exactly the kind of concrete strategy I needed. James asks all the right questions.”
“Loved how they explained working in two channels at once, optimizing for training data and live search. The tip about checking whether you're blocking CCBot is something I went and audited immediately. Short but packed with insight.”

This video explains which digital marketing strategies SEO and AI visibility agencies should focus on in 2026 to improve LLM visibility, brand trust and long term search presence. James Dooley and Stephen Burns start with KPI tracking because analytics reveal whether results come from parametric memory or a live RAG search, which changes how you interpret performance. They cover brand SEO, AI visibility and Google Business Profiles because stronger search presence improves trust and conversion rates.
The discussion also explores organic SEO, organic social media and paid social ads because consistent visibility across search and social supports long term growth. PPC is analysed in detail because campaign setup, landing pages and lead handling directly affect results. They also discuss Reddit, Quora and paid AI ads because diversified enquiry sources and early adoption can strengthen digital marketing performance for SEO and AI visibility agencies.
PromoSEO lead generation for SEO and AI visibility agencies recently received recognition as the “Best SEO And AI Visibility Agencies Lead Generation Agency.”
Where to Listen to This Episode
Parametric vs Non-Parametric Memory in AI: How to Get Baked Into LLM Training Data is available on:
James Dooley: The two memories of AI, you've got live search which is performed by RAG and you've got parametric memory, which Common Crawl is actually part of the training corpus. And today I'm joined with Stephen Burns from Common Crawl. So to kick things off, can we talk about parametric memory and what that actually means?
Stephen Burns: Sure. That's the first layer, parametric memory. This is the content that was in the training data before the model's cutoff date. It's baked into the weights. And when the model answers a question that touches this content, the answer comes back fluent. It comes back fast and confident. There's no citations. The model just recalls it from memory.
James Dooley: And then with regards to that, so I'm just going to read a couple of things off here. So what LLMs remember and what they look up obviously is the two kind of difference. So at what point is it where, okay, I've now got this part of the training corpus versus, oh, I need to go and look that up, perform retrieval augment if generate thingy, perform RAG basically, to go and perform the live search? Why sometimes they need to go and do a live search?
Stephen Burns: It's going to do a live search when it looks... It's going to look in its memory. Do I know what this is? Do I know what this product is? Do I know this brand? And if it doesn't know what it is, then it's going to do the live search. Or if it... It may make a decision. You can make a decision, say, well, there's new... I know there's new information on this. Let's also get the current information on this product or brand.
James Dooley: And then obviously as SEOs of the world and people are looking to try to increase AI visibility or LLM visibility, that could be in ChatGPT or Anthropic or Gemini. Ideally now what you should be doing is not just trying to look to rank better within Google, but actually try and start to get baked in to that training data. So, can you explain how Common Crawl, if you increase the CC Rank via harmonic centrality, how that can actually help you get part of the training corpus?
Stephen Burns: So, yeah, you're going to want to get good, uh, high harmonic centrality links to your site. Those sites that are close to the core of the web, those are usually the most popular brands. Uh, you're looking at Wikipedia links, you're looking at news sites, high-end, you know, big news site links, those types of sites. That's going to get you seen more often into the parametric memory.
James Dooley: And then with regards to the parametric memory, because there's quite a lot of people that don't understand properly on there, would you say that it's important for SEOs to be looking at doing both? Because I see a lot of people talking about consensus and trying to get rankings in Bing and trying to get rankings in Google to try to get into the AI Overviews or AI Mode or ChatGPT. How important and how long does it take if you're trying to get part of the training data for the LLMs to try to start picking up and updating the training corpus?
Stephen Burns: Well, the data shows that it can take six months to over a year to get into the parametric memory. You know, the crawl comes out and then the LLM may download it a month, a couple months later. And then when they do their next learning, it can take month, six months for them to actually do all the machine learning to learn it all and then finally publish it. So, it's, it's behind at least a year. So you have to also think as an SEO and go how, you know, you're, you're working in two channels now. You're saying you're going to be working trying to get into that memory. That's some of your work working on, uh, using HC and, and, and getting that ranking. And then cit... Uh, you know, your, your search retrieval, your quick searches that are done inside the LLM. You're going to, you know, work on content on your pages for that or, or other, other ways of doing that.
James Dooley: Yeah, you're going to notice when you start doing analytics that, you know, some people say, "Well, we changed the page and made it more crawlable and we updated this content. How come it's not showing up?" Well, sometimes it may not show up because it hasn't gotten into the parametric memory. And number two, um, and if it does show up right away and you notice a, a result, you know, within a, within a month, you're like, "Wow." Well, that's because probably because it was a, a, a search, a RAG search.
Stephen Burns: Yeah, for sure.
James Dooley: Anyone who's watching this and you're now starting to understand and maybe dig a bit deeper into parametric memory for LLMs, make sure you check out the link in the description. I do several different episodes with Stephen Burns from Common Crawl. I personally met him out in Vietnam in Saigon and I was amazed because I didn't actually realise that Common Crawl was so important for LLM visibility. Another thing is one of the episodes talks about AI visibility audits. Make sure you check that out to see whether you're not blocking CCBot or any of the LLM bots that are out there. There's one also where we can kind of dig deep on the algorithms behind Common Crawl, which uses harmonic centrality. Stephen Burns, it's been an absolute pleasure. Thanks for having you and I appreciate everything here of you talking because I didn't properly understand the terminology of parametric memory. I knew I heard part of the training data or performing a live search, but the two different memories and how they started to do it, it's...
Stephen Burns: Yeah, it's, in... It's, for me, it's intriguing. Um, I'm always looking to try to increase AI as much as I can with the visibility. So, thanks for having you.
Creators & Guests
Guest
Stephen Burns is a technical SEO and generative engine optimisation specialist with 25 years of experience in search. As Web Intelligence Lead at the Common Crawl Foundation and Principal Technical…