What RAG is, in plain terms
The system first searches your documents for the passages that bear on a question, then asks the AI to write an answer using only those passages and to cite them. That is the difference from a general chatbot, which answers from what it learned on the open internet and knows nothing about your pricing, policies or exceptions.
The cost drivers
| Driver | Why it moves the price | What to ask a vendor |
|---|---|---|
| Number and type of sources | Drives, CRM notes, tickets, email, PDFs and wikis each need their own connection and cleaning | Which sources will you connect, and who maintains the connections? |
| Quality of the content | Duplicates, old versions and scans need cleaning before search works well | How do you decide which version of a document wins? |
| Permissions | If the finance folder must stay invisible to the sales team, retrieval has to respect that | How is access enforced at question time? |
| Testing | A set of real questions with known good answers is the only way to know it works | How do you measure accuracy, and can I see the results? |
| Where it runs | A hosted model, your private cloud or your own hardware have different costs and different risks | Does our data ever leave our control? |
| Integration | A search box is simple; answers inside the CRM, Slack or the phone system are more work | Where will my team actually use it? |
What a build includes
- An inventory of sources and who may see what.
- Cleaning, structuring and tagging the material, with dates and versions.
- Retrieval tuned to your content, and answers written only from what was found, with the source shown.
- A test set of real questions, scored before launch and after every change.
- Access rules by role, a way to flag a wrong answer, and documentation your team can run.
Running costs to plan for
- Hosting, and either model usage fees or the hardware to run a model yourself.
- Re-indexing as documents change, so answers stay current.
- Monitoring and a regular look at the questions the system could not answer.
When RAG is the wrong tool
- You have a handful of documents: pasting them into a single prompt may be enough.
- You need the system to take actions, not just answer: that is an AI agent.
- You need the model to change how it writes or classifies: that points to fine-tuning.
How we scope one
We look at the material and the questions first, then write the scope and price down before any build starts. If you are not sure whether RAG is the right fit, the operations audit answers that along with everything else.