Your RAG pipeline might be lying to you.
In this episode, Yuval and Niv from AI21 break down why most Retrieval-Augmented Generation (RAG) systems pass benchmarks but fail in the real world.
We’re talking:
- Why popular RAG benchmarks reward the wrong things (and hide real failures)
- The chunking trap: how bad segmentation can tank even good retrieval
- When LLMs nail the answer—but your pipeline still screws it up
- Structured RAG: the fix for aggregative data like financial reports
- Real-world evaluation tips (so you stop shipping broken systems)
- Bonus: What do Seinfeld trivia, World Cup stats, and enterprise SharePoint have in common? Your RAG pipeline probably chokes on all of them.
Speakers

Niv Granot
Tech Group Lead @ AI21

Yuval Belfer
Sr. Developer Advocate @ AI21


