HB
Back to articles
RAG & AI Architecture4 min read

Fixing legal RAG retrieval with Azure AI Search

Preserving tables, combining retrieval methods, and checking failed answers

By Hasan Butt · Updated

The retrieval failure came before generation

In my legal RAG project, schedule tables and nested clauses were being merged into surrounding text. The retrieved passage could mention the right topic while omitting the number or exception needed to answer the question. Adding more passages increased context size without reliably restoring the missing structure.

The project describes a corpus of 3,145 documents and a 200-question validation suite. It reports a faithfulness score rising from 0.72 to 0.98. The evaluation artifacts are not published with this guide, so those are author-reported project results, not a reproducible benchmark or a guarantee for another corpus.

Preserve the table and its surrounding rules

Keep headings, table rows, footnotes, and source-page references together when they are needed to interpret a value. Store the document identifier and section path with each chunk. A smaller search chunk can point to a larger parent section, but the parent must remain within the generation budget.

For example, a test question might ask which fee applies on a particular date. This is an illustrative test pattern, not a client question. Check that retrieval returns the fee row, effective date, and relevant exception together. A passage containing only the fee name should fail the test even if its similarity score is high.

Combine lexical and vector retrieval

Lexical retrieval is useful for exact section numbers and domain vocabulary. Vector retrieval can help when the question uses different wording from the source. Azure AI Search can run both in a hybrid request and combine their rankings using reciprocal rank fusion.

The following asynchronous example requires the Azure Search Python SDK, an existing index with a vector field named embedding_vector, and a query embedding with the dimensions configured for that field. Supply an authenticated async client. It is an integration example, not an executed benchmark.

from azure.search.documents.aio import SearchClient
from azure.search.documents.models import VectorizedQuery

async def retrieve(
    client: SearchClient,
    query: str,
    embedding: list[float],
    authorized_filter: str,
):
    results = await client.search(
        search_text=query,
        vector_queries=[VectorizedQuery(
            vector=embedding,
            k_nearest_neighbors=50,
            fields="embedding_vector",
        )],
        filter=authorized_filter,
        top=20,
    )
    return [document async for document in results]

Build authorization filters from trusted server-side identity and access rules. A model may suggest a date or jurisdiction filter, but it must not decide which tenant's documents a user is allowed to read.

Rerank only when the evaluation supports it

Reciprocal rank fusion combines ranking positions rather than comparing incompatible raw scores. Each candidate receives a contribution of 1 / (k + rank) from each ranking where it appears. Azure performs fusion for hybrid queries; the application does not need to repeat it.

A reranker can then compare the question with the candidate passages more closely. Azure semantic ranking or a separate service such as Cohere adds another cost and latency step. Compare retrieval alone with retrieval plus reranking on the same questions before making it mandatory. Pick the candidate count and final context size from those results.

Make unsupported answers visible

Ask the generator to attach source identifiers to factual claims, then check that those identifiers exist in the retrieved context. A valid identifier alone does not prove the source supports the sentence. Review the cited passage and include unsupported-answer cases in evaluation.

For known statutory tables, explicit lookup rules can be useful. Their limits matter: the rule needs to select the correct jurisdiction, version, and effective date. When those inputs are missing, ask for clarification or report insufficient evidence instead of guessing.

Keep a failure set, not just an average score

Include exact-number questions, nested exceptions, missing sections, conflicting versions, and questions the corpus cannot answer. Separate retrieval failures from generation failures. Keep the judge model, prompts, dataset revision, and scoring method with each evaluation run.

Inspect latency and token use alongside answer quality. A change can improve the mean score while making important questions worse. Keep those regressions visible and rerun the same set after parser, embedding, prompt, or model changes.

More engineering notes