AI & LLM Engineering65 min total · 11 parts
AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling
Part 8 of 11 · ~3 min
Hallucination, and Why RAG Doesn't Cure It
It's tempting to treat retrieval as the fix for made-up answers, and it genuinely helps — but Owen's refund question, revisited one more time, shows exactly where it stops helping. Suppose the retrieved chunk is "Cancelling Your Subscription," which is the right document, correctly retrieved, and it covers cancellation thoroughly without ever using the word "refund" at all, because there simply isn't a refund policy to describe. A model asked "can I get a refund" with that chunk sitting in front of it will sometimes still answer with something like "Yes, refunds are available within 14 days of cancellation" — a plausible-sounding SaaS refund policy pulled from its training, laid neatly on top of a context that never said any such thing.
This is hallucination, and the RAG version of it is more dangerous than the bare-prompt version from chapter one specifically because it looks more trustworthy — there's a citation number sitting right next to the fabricated claim, and a citation reads as proof, when all it actually proves is that the model looked at that document, not that the document supports what got written next to its number. The retrieved chunk did its job. The failure is that the model, faced with a question its context doesn't actually answer, filled the gap with its training-data instincts about SaaS refund policies in general rather than reporting the gap honestly.
Two mechanisms cut this down, neither of which is a complete guarantee, and it's worth being honest about that rather than overselling either one.
Explicit permission to say "not in the sources." The system prompt from chapter four already says this, but it's worth restating why the wording matters: "answer the question" is an instruction that implicitly rewards producing an answer, any answer, whereas "if the sources don't address this, say so" gives the model a legitimate, named, non-failure path that isn't "give up" — and a model, like most systems optimized to be helpful, will take the productive-looking path over the honest-looking one unless the honest one is spelled out as acceptable, even expected.
A grounding check on the way out, catching what the prompt alone didn't:
function citesRealSource(reply, matches) {
const cited = [...reply.content.matchAll(/\[(\d+)\]/g)].map((m) => Number(m[1]));
if (cited.length === 0) return false; // no citation at all — treat as ungrounded
return cited.every((n) => n >= 1 && n <= matches.length); // every citation points at a real source
}
async function askLoopGrounded(question) {
const matches = await retrieveContext(question);
const reply = await askLoopRAG(question);
if (!citesRealSource(reply, matches)) {
return { content: "I don't have a confident answer to that from our help center. Let me connect you with a human who can check.", escalate: true };
}
return reply;
}
That check is deliberately cheap and deliberately shallow — it confirms the model referenced some real source, not that the cited source actually supports every claim in the sentence next to it, which would mean re-verifying the claim against the source text, a meaningfully harder problem. It catches the reply that cites nothing at all, which is a strong, cheap signal that the model reached past its context. It does not catch a reply that cites [1] correctly for the cancellation date and then quietly adds an invented refund clause in the same sentence. That gap is exactly why the evaluation chapter near the end treats a hand-built golden question set, checked regularly against real answers, as a permanent part of running this feature — not a one-time pre-launch task.