Syntax Station

Insights / AI & Agents

Enterprise RAG That Works: Hybrid Search, Reranking and Honest Citations

Why basic vector search breaks on real company documents, and the retrieval pipeline we use to build AI assistants that find the right answer and show where it came from.

By Syntax Station Engineering · · 4 min read

Key takeaways

  • Pure vector search misses exact terms like product codes, clause numbers and names. Combine it with keyword search.
  • A reranking step that reads the top results before they reach the model is one of the cheapest quality wins available.
  • How you split documents matters as much as which model you use. Keep headings, tables and parent sections attached to each chunk.
  • Make the assistant cite its sources and say "I don't know" when retrieval comes back empty.

Building a demo that answers questions from a PDF takes an afternoon. Building an assistant that a legal team, a support desk or an engineering organization trusts every day is a different job. The gap between the two is almost entirely about retrieval.

Why the basic approach fails

The tutorial version of RAG works like this: split documents into fixed-size chunks, turn each chunk into an embedding, store them in a vector database, and at question time fetch the five closest chunks.

On real company data, this breaks in predictable ways:

  • Exact terms get lost. Embeddings capture meaning, not spelling. A question about "SKU 4471-B" or "clause 14.3" can return passages that are about similar topics but never mention the exact item.
  • Context gets cut in half. A fixed 500-token window can split a table from its header or a policy from the exception two paragraphs later.
  • Old and new versions compete. Without metadata, a 2023 policy and its 2026 replacement look equally relevant.
  • Permissions are ignored. If retrieval does not filter by who is asking, the assistant can quote documents the user should never see.

The retrieval pipeline we use

1. Structure-aware chunking

We parse documents with their structure intact: headings, lists, tables and page numbers. Each chunk carries its section path ("HR Policy > Leave > Parental leave") and links to its parent section. When a small chunk matches, the model can receive the surrounding section too.

Every query runs two searches in parallel: a keyword search (BM25) that is excellent at exact terms, and a vector search that is excellent at meaning. The results are merged, typically with reciprocal rank fusion. This single change fixes most "it can't find the obvious document" complaints.

3. Metadata filters

Document type, department, region, effective date and access group are stored with every chunk and applied as filters before ranking. Users only retrieve what they are allowed to read, and current documents win over archived ones.

4. Reranking

The top 30 to 50 candidates go through a reranker, a model that reads the question and each passage together and scores true relevance. Only the best handful reach the language model. Reranking is cheap, fast and consistently one of the largest quality improvements in the pipeline.

5. Grounded answers with citations

The prompt requires the model to answer only from the supplied passages and to cite them. The interface shows those citations as links, so users can check the source in one click. When retrieval returns nothing useful, the assistant says so instead of guessing.

Measuring quality

You cannot improve what you do not measure. Before launch we build a test set of 100 to 300 real questions with known answers and track two numbers separately:

MetricWhat it tells you
Retrieval recallDid the right passage appear in the results at all?
Answer faithfulnessDid the answer stick to what the passages say?

Separating the two shows where to work. Low recall means fixing chunking, search or filters. Low faithfulness means fixing the prompt or the model.

Keeping it current

Documents change. A production RAG system needs an ingestion pipeline that picks up new and edited files automatically, removes deleted ones and re-indexes on a schedule. It also needs logging of every question and answer (with care for personal data) so you can see what people ask and where the assistant falls short.

When RAG is the right choice

RAG fits when answers live in documents that change, when users need to verify sources, and when different people may see different data. If you are deciding between approaches, our comparison of fine-tuning, RAG and prompt engineering walks through the trade-offs.

Frequently asked questions

What is RAG in AI?

Retrieval-augmented generation (RAG) is a pattern where an AI system first searches your documents for relevant passages, then gives those passages to a language model so its answer is grounded in your data instead of only its training.

Is RAG better than fine-tuning?

For answering questions from company knowledge, usually yes. RAG keeps answers current when documents change, can cite sources and respects access permissions. Fine-tuning is better for changing a model's style or teaching it a narrow task format.

Which vector database should I use?

For most companies, the database you already run is fine: PostgreSQL with pgvector handles millions of chunks comfortably. Dedicated vector databases make sense at very large scale or when you need advanced filtering and multi-tenant isolation out of the box.

Related reading