Retrieval-augmented generation (RAG)

Retrieval-augmented generation (RAG) is a technique in which relevant passages are fetched from a document store at question time and given to a language model. Its answer is then based on those documents rather than on memory alone.

Also called: retrieval augmentation, knowledge retrieval

A business has more information than fits in a prompt, service descriptions, policies, FAQs, price lists, manuals. RAG indexes it, and when a caller asks something, the system searches the index for the most relevant passages and hands them to the model alongside the question. The model then answers from what it was just shown.

The benefit is that the knowledge base can be large and can change without retraining anything: update the document, and the next call reflects it. The limitation is that retrieval has to find the right passage; a question phrased unusually, or a fact buried in a table, can be missed, and the model then answers without it.

RAG is grounding for documents. It is different from real-time integration, which fetches live records from a system of record, a calendar slot, an order, through function calling rather than from a text index.