Retrieval-augmented generation, usually called RAG, connects a language model to information it can use as evidence for an answer. For a business, that might include product documentation, approved policies, support knowledge, or internal operating procedures. The difficult part is not making a model summarize a document once. It is making the right information available to the right person, at the right time, with a result that can be checked.
Consider an illustrative service company whose staff search several systems to answer questions about customer onboarding. A RAG feature could reduce that search effort. It could also expose a restricted document, quote an obsolete policy, or confidently combine two instructions that apply to different customer groups.
This guide explains how to move from a demonstration to a maintainable knowledge product. It covers source ownership, ingestion, retrieval, permissions, citations, evaluation, and operations. The recommendations are an engineering approach to the problem, not a claim that any particular retrieval architecture guarantees correct answers.
Key takeaways
- Treat source quality, access control, and freshness as product requirements.
- Evaluate whether the right evidence was retrieved separately from whether the answer used it correctly.
- Preserve document identity, version, permissions, and citations through the pipeline.
- Start with one valuable knowledge workflow and expand after measuring real task quality.
Choose the knowledge task and its boundaries
Begin with the questions people need to answer and the decisions those answers support. Searching setup instructions is different from interpreting a contract or deciding eligibility for a benefit. The consequence of an incorrect answer should influence the scope and review process.
Collect representative questions from actual work with sensitive information removed where appropriate. Include the vocabulary people use, the context they assume, and the follow-up questions they usually ask. A system evaluated only on questions copied from document headings will appear more capable than it is.
Define which sources are authoritative for the task. A current policy may override an old support conversation, while a customer-specific agreement may override a general guide for authorized users. Write those precedence rules explicitly rather than expecting the model to infer them reliably.
Specify what the feature should do when the answer is absent or ambiguous. It may ask for the customer group, show relevant documents without a conclusion, or route the question to an owner. A useful knowledge product does not need to answer every question.
Inventory and repair the source material
List the systems containing relevant knowledge and identify an owner for each source. Record document types, update frequency, access rules, and known quality issues. Duplicated pages, contradictory instructions, and scanned documents require different preparation work.
Separate authoritative material from drafts and informal discussion. A message that solved one unusual case should not silently become general policy. Mark the scope of each source, including product version, location, customer type, or effective period where relevant.
Improve content where the source is genuinely unclear. Retrieval tuning cannot fully compensate for a document that never states the prerequisite or exception users need. A small editorial effort can sometimes improve answers more than another model or search layer.
Create a content ownership process. When the product changes, someone must update the relevant knowledge and retire obsolete instructions. Without that process, the system will gradually become a convenient way to retrieve outdated information faster.
Build ingestion around traceability
Every indexed passage should remain connected to its original document, location, version, and permission context. Preserve enough information to show a useful citation and to remove or update the passage when the source changes. A text fragment without provenance is difficult to trust or maintain.
Extract content in a way that preserves meaning. Tables, headings, lists, and captions can contain relationships that are lost when a document is flattened into plain text. Test representative documents and inspect the extracted result rather than assuming a parser handled them correctly.
Split material into retrieval units that match the content. A short procedure may need its prerequisites and steps together, while a long reference manual may benefit from smaller sections. Avoid choosing one chunk size solely because it worked in a tutorial.
Track ingestion status and failures. A connector that silently stops updating can leave the product serving old information while appearing healthy. Provide an operational view of source coverage, last successful updates, rejected documents, and the reasons for those rejections.
Compare retrieval approaches on real questions
Microsoft’s RAG overview describes grounding answers in proprietary content and discusses keyword, vector, and hybrid retrieval approaches. These are options to evaluate, not a guarantee that adding more retrieval stages will improve every task.
Build a baseline with a simple search approach and a known question set. Check whether the relevant evidence appears in the returned results and whether irrelevant material crowds it out. Exact product identifiers, abbreviations, and policy names may behave differently from conversational paraphrases.
Test filters for scope such as product version, organization, and effective date. Better filtering can be more valuable than retrieving a larger number of passages. Too much loosely related context can make the answer harder to ground and increase latency and cost.
Add reranking or query rewriting only when evaluation shows a useful improvement. Inspect failures to understand whether the problem is vocabulary, missing content, poor extraction, scope, or ranking. Each has a different remedy. Treat retrieval as an information-quality problem rather than a contest to assemble the most elaborate pipeline.
Enforce permissions before evidence reaches the model
A user should not receive restricted content simply because the final answer is expected to omit it. Apply access controls during retrieval and validate the scope on the server. Once sensitive material is in model context, a prompt instruction is not an adequate substitute for authorization.
Decide how source permissions are represented and refreshed. Group membership, document access, and organization context can change. A stale permission index may continue exposing content after the source system revokes access, so define and test the propagation behavior.
Partition caches according to the authorization context. An answer generated for an administrator must not become a shared cached response for an ordinary user. Include relevant source versions and scope in cache design, and avoid caching private answers more broadly than intended.
Test cross-organization access, revoked membership, and mixed-permission result sets. Use fixtures that make leakage detectable. Permission testing belongs alongside answer-quality evaluation, because a highly accurate answer can still be unacceptable if the user was not entitled to its evidence.
Make answers verifiable and uncertainty useful
Ask the model to answer the specific question using the available evidence and to distinguish supported facts from missing information. Keep instructions clear, but verify behavior through evaluation rather than assuming the prompt guarantees it.
Attach citations to the claims they support. A list of links at the end can look reassuring while failing to show which source supports which statement. Let users open the relevant document or section, and preserve the title and version where those details matter.
Avoid presenting a similarity score as a probability that the answer is correct. Retrieval relevance and factual correctness are different properties. If the system lacks enough evidence, explain the gap in plain language and offer a useful next step.
Handle conflicting sources explicitly. The application can apply documented precedence rules or show the conflict for review. It should not silently combine incompatible instructions into a new policy. In the onboarding example, a general guide and a customer-specific exception should remain distinguishable.
Evaluate retrieval and generation separately
Create a set of questions with expected supporting documents and reviewed answer requirements. For retrieval, measure whether the necessary evidence is available. For generation, assess whether the answer is accurate, complete enough, properly scoped, and supported by that evidence.
Include questions whose answer is not in the corpus. The system should recognize the limitation rather than invent a plausible policy. Add ambiguous wording, outdated terminology, conflicting documents, and multi-step questions that require more than one source.
Use deterministic checks where possible, such as required identifiers or forbidden data exposure, and human review for nuanced interpretation. Model-based grading can help scale review, but calibrate it against examples people have judged and inspect disagreements.
Keep a stable evaluation set for comparison and a growing set of production failures. Run both when changing extraction, chunking, search settings, prompts, models, or source rules. A retrieval improvement on one question type may harm another, so examine the distribution rather than only an overall score.
Design freshness, deletion, and version changes
Define how quickly different sources need to update. A product release note and a live account entitlement have different freshness requirements. Some questions may be better answered through a current API than a periodically indexed document.
Propagate deletion and access changes as deliberately as new content. Track which indexed passages came from each source so they can be removed. Clear or invalidate affected caches, and verify that a deleted document no longer appears in answers or citations.
Make version transitions observable. If an ingestion job partially fails, operators should know which sources remain old. Consider whether the system should continue serving the previous version, restrict certain questions, or show a freshness notice according to the task’s consequences.
Test an end-to-end update: change a source, wait for the expected processing, ask a relevant question, and inspect the returned evidence. Do the same for deletion and permission revocation. These checks reveal whether the operational promise matches the actual pipeline.
Control latency, cost, and failure behavior
Measure the stages separately: source access, retrieval, reranking, model generation, and any additional tool calls. A slow answer may come from an external connector rather than the model. Stage-level measurements help target improvements without weakening quality blindly.
Limit retrieved context to what the task needs and set budgets for query expansion or repeated searches. A feature that decomposes every simple question into many searches can become expensive and slow without improving the answer.
Provide a useful fallback. If generation is unavailable, the product may still show relevant search results. If a particular source is unavailable, explain the limitation rather than implying a complete search was performed. Preserve enough diagnostics for the team to investigate the failure.
Track cost per successfully resolved knowledge task. Include failed requests and retries. Compare the result with the previous search or support workflow, including review effort. The business value depends on useful answers and saved work, not merely a low cost per model call.
Launch with one accountable knowledge domain
Choose a domain with a clear owner, manageable source set, and real user demand. Review the sources, establish a baseline, and pilot the feature with people who perform the task. Give them a simple way to report a wrong answer or missing document.
Inspect feedback at the stage where the problem occurred. A wrong answer may need better source content, retrieval filters, generation instructions, or interface design. Avoid treating every issue as a prompt problem.
Expand only after the team can operate updates, permissions, evaluation, and support reliably. Each new knowledge domain introduces vocabulary, source quality, and ownership questions. A production RAG system grows sustainably when those responsibilities scale with its coverage, so users can understand what it knows, why it answered, and when they should seek another source.
Frequently asked questions
Is RAG the same as fine-tuning?
No. RAG supplies retrieved evidence at answer time, while fine-tuning changes model behavior through additional training. They address different needs and can sometimes be combined. For changing business knowledge with source citations, retrieval is often a useful starting point, but its quality and access controls still need careful design.
Do we need a vector database?
Not automatically. Start with the retrieval requirements and evaluate available search capabilities. Keyword, vector, and hybrid approaches can perform differently across identifiers, terminology, and conversational questions. An existing search service may be sufficient. Choose based on measured evidence quality and operational fit.
How much content should we index first?
Enough to cover one useful, well-defined workflow with accountable owners. A smaller reviewed corpus can be more valuable than a large uncurated collection. Expand after you can measure retrieval, answer quality, freshness, and permission behavior. Volume alone does not make the system knowledgeable.
Can a RAG assistant answer confidential employee questions?
Potentially, but only with appropriate source permissions, identity handling, data arrangements, and review of the task’s consequences. Do not retrieve restricted material and rely on the model to hide it. Test revocation, organization boundaries, and caches, and provide a clear route for questions requiring professional or organizational judgment.
Working through this in your product?
We can help you turn these decisions into a practical plan and working software.
Knowledge search and AI integrationTalk to the team
