TL;DR

  • Traditional RAG falls short on complex enterprise queries, leading to incomplete or unreliable answers.
  • Structured RAG (S-RAG) transforms unstructured documents into structured, query-aware representation for precise, auditable reasoning.
  • AI21 Maestro’s hybrid architecture blends structured and embedder-based retrieval, improving accuracy by up to 60% with near-perfect recall.
  • Tailored for enterprise scale: automatically infers or lets users define schemas to handle millions of documents with full transparency and control.
  • Result: reliable, explainable answers you can trust in compliance, reporting, and mission-critical workflows.

We’ve spent years working with enterprises trying to turn their data into something usable. Every team had the same hope: RAG would finally unlock their knowledge for scaled use. But by the time we got involved, reality had already hit. The answers were fluent, yes, but they weren’t always reliable.

Let’s say you ask about performance trends or supplier comparisons, and the LLM-based system gives you only part of the picture. It might surface a few relevant excerpts but miss key data points, or summarize loosely connected information that doesn’t directly answer the question. In finance and compliance, those gaps aren’t harmless, they’re major risks.

We knew RAG wasn’t totally broken, but it needed help to become reliable enough for enterprise use. So we built Structured RAG (S-RAG), an enhancement that brings structure and reasoning to retrieval for precise, auditable answers. It’s built into AI21 Maestro, our AI orchestration system powering enterprise agents for mission-critical tasks.

Want to get more technical? Deep dive into our Structured RAG research paper.

Unlike traditional RAG, which retrieves text chunks and hopes the LLM answers correctly based on limited context, S-RAG extracts information offline from unstructured documents into a relational database. At runtime, it uses SQL queries to retrieve and reason over that structured data. This means analytical questions that require aggregation, comparison, or filtering — the kind enterprises rely on — can finally be answered accurately and transparently.

Example: When asked a simple financial question that most RAG systems fail to answer correctly, Maestro’s S-RAG retrieves the precise value directly from structured data rather than guessing based on text proximity.

Here’s what’s happening behind the scenes:

SELECT 
    “current_liabilities” / 1000000 AS current_liabilities_millions
FROM 
    “SEC_Report”
WHERE 
    LOWER(“company_name”) = 'netflix' 
    AND “fiscal_year” = 2017

Where embedder-based RAG falls short

Embedder-based RAG, the foundation of the first generation of knowledge agents, marked an important milestone in the evolution of retrieval-augmented generation. By embedding queries and documents into high-dimensional vectors, it aimed to retrieve semantically “close” text snippets and feed them to an LLM for reasoning.

But for enterprises, this approach falls short in three critical cases.

1. Answering aggregative questions

We’ve seen customers, especially in finance and compliance, hit walls with seemingly simple analytical queries:

  • “What was the maximum capital expenditure across all subsidiaries last year?”
  • “Who are the top five suppliers by on-time delivery rates?”
  • “How have safety incidents trended quarter over quarter?”

These are not “needle-in-a-haystack” lookups, where a few chunks of text hold the answer. To correctly answer these analytical queries, the system has to filter, compare, and aggregate data points across potentially dozens or hundreds of records.

Embedder-based RAG has no generalized way to perform these operations. It retrieves a predefined number of chunks of text and passes them to an LLM, which must attempt reasoning inside the limited context window using very limited arithmetic capabilities. 

2. Questions that require exhaustive coverage

Enterprise teams often ask questions that demand complete, exhaustive lists:

  • “List all contracts expiring before 2025 with penalty clauses over $1M.”
  • “Which employees have certifications that will expire this year?”
  • “Show all markets where regulatory changes affect reporting requirements.”

Missing even one item is not an inconvenience, but a compliance risk. Yet embedder-based retrieval is inherently probabilistic and optimized to find the single best matching chunk rather than all relevant evidence. Because it only fetches a subset of documents based on similarity scoring, there can never be guarantees of full retrieval. This is unacceptable in domains like financial reporting or regulatory compliance, where exhaustive coverage is as important as correctness.

3. Dense or non-indicative corpora

Finally, consider corpora like technical documentation or regulatory filings. These are dense, repetitive, and attribute-driven. Documents may differ only by a few numbers or clauses, and queries may focus on precise attributes like clause numbers or model IDs. In these settings, embedding similarity collapses. Retrieval becomes noisy because embedder models haven’t learned the underlying semantics of those terms, so irrelevant but “close” passages often crowd out the real answer.

For example, in financial analysis queries, most reports look nearly identical and contain similar terminology — “total liabilities,” “shareholder equity,” “net income”. To an embedder, they all seem relevant. As a result, semantic retrieval often surfaces a stack of irrelevant documents that appear contextually close but don’t actually answer the question.

Enter Structured RAG.

Enhancing RAG with AI21 Maestro

AI21 Maestro solves these problems with Structured RAG, by expanding the range of queries RAG systems can accurately process in enterprise environments.

Instead of treating documents solely as unstructured text, AI21 Maestro leverages structure at ingestion to complement the drawbacks of embedder-based retrieval. It analyzes documents to detect recurring patterns. For example:

  • In financial filings, every report describes attributes like revenue, operating expenses, and capital expenditure.
  • In HR resumes, every record contains education, years of experience, and certifications.
  • In contracts, clauses follow standard sections like termination, penalties, and governing law.

But because documents are written independently, the same value can appear in many different formats — for example, “1,000,000,” “1M,” or simply “1.” Without standardization, reasoning across such documents is nearly impossible.

AI21 Maestro automatically infers a schema that captures these attributes. With structured RAG, each document is transformed into a structured record aligned to this schema, with consistent formatting, terminology, and validation. Users can review, edit, and adjust the automatically generated schema to fit their specific needs before ingestion. The entire corpus becomes a structured database, with links to the original text for transparency and control.

At inference time, when a user asks a relevant question that relates to one of the available schemas, AI21 Maestro translates the natural-language question into a formal SQL query over this structured database.

Structured RAG AI21 MAestro

This enables precise analytical operations that traditional RAG can’t perform — max/min calculations, trend analysis, complex filters — while maintaining traceability.

  • “What’s the highest ARR across subsidiaries?”
  • “Show average headcount growth by year.”
  • “Which contracts signed after 2020 include arbitration clauses?”

After executing the SQL, Maestro returns the result to the LLM for a fluent, business-ready answer.

Performance and hybrid retrieval

This approach produces up to 60% higher accuracy on aggregative queries than embedder-based systems and a near-perfect recall for exhaustive coverage questions (given the right schema).

Answer comparison on aggregative evaluation sets

A comparison of embedder-based RAG (with gpt-o3 as generational model), OpenAI Responses API, and Structured RAG on aggregative-question benchmarks.

Answer recall on aggregative evaluation sets

But not all information fits neatly into a schema. Some attributes are rare, or a question may require reasoning over narrative text. In those cases, AI21 Maestro uses Structured RAG to narrow the dataset first, then applies embedder-based retrieval on that reduced set.

This hybrid retrieval approach balances exhaustive coverage with flexibility, ensuring that no relevant information is missed while still supporting nuanced, text-heavy questions. It also sidesteps the failure modes of dense corpora by querying exact attributes first.

As shown below, Maestro’s hybrid approach clearly outperforms embedder-only RAG and competitive RAG systems across real-world benchmarks.

Comparison of Maestro Hybrid RAG, a classic embedder-based RAG and OpenAI Responses API over FinanceBench, a large financial analyst-style benchmark.

From answers to decisions

For enterprise leaders, the implications are clear.

  • Trust the answers: Structured reasoning delivers outputs that are accurate, auditable, and defensible in high-stakes settings. AI21 Maestro’s dynamic hybrid architecture combines structured and embedder-based retrieval to get the best of both worlds, improving results while keeping the classic pipeline as a reliable fallback.
  • Scale without fear: AI21 Maestro handles millions of documents, dense technical corpora, and ever-evolving regulatory requirements. S-RAG adapts to your specific queries — automatically inferring schemas based on query context or letting you define them yourself for full control.

Because Structured RAG operates within AI21 Maestro’s dynamic hybrid pipeline, every answer you get is more accurate than the baseline RAG results you already trust.

By addressing the blind spots of embedder-based approaches, AI21 Maestro turns unstructured enterprise data into a structured decision-making engine. For enterprise, that means answers you can act on, workflows you can automate, and risks you can mitigate.

Hear the truth behind why RAG isn’t solved yet — and how Structured RAG changes that.

🎧 YAAP Podcast: “RAG Is Not Solved – Your Evaluation Just Sucks”

🎥 AI Engineer World’s Fair: A deep dive into building reliable retrieval systems.

Structured RAG is now part of AI21 Maestro. Contact us to learn how it retrieves better answers you can trust.