Your RAG pipeline might be lying to you.
In this episode, Yuval and Niv from AI21 break down why most Retrieval-Augmented Generation (RAG) systems pass benchmarks but fail in the real world.

We’re talking:

  • Why popular RAG benchmarks reward the wrong things (and hide real failures)
  • The chunking trap: how bad segmentation can tank even good retrieval
  • When LLMs nail the answer—but your pipeline still screws it up
  • Structured RAG: the fix for aggregative data like financial reports
  • Real-world evaluation tips (so you stop shipping broken systems)
  • Bonus: What do Seinfeld trivia, World Cup stats, and enterprise SharePoint have in common? Your RAG pipeline probably chokes on all of them.

Speakers

Niv Granot - Algorithms Group Lead @ AI21

Niv Granot
Tech Group Lead @ AI21

Yuval Belfer - Sr. Developer Advocate @ AI21

Yuval Belfer
Sr. Developer Advocate @ AI21