AI & LLM Engineering65 min total · 11 parts
AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling
Part 5 of 11 · ~2 min
Retrieval-Augmented Generation
RAG is the name for the whole pattern: instead of asking a model to answer from what it already learned during training, you retrieve the actually-relevant material at the moment of the question and hand it over as part of the prompt, so the model is generating from what's in front of it rather than from a vague, dated, sometimes-wrong memory of "SaaS products in general."
There are two separate phases, running on entirely different schedules, and mixing them up is a common source of confusion. Ingestion happens offline, once per document (and again whenever a document changes) — embedding every help article and storing the result. Retrieval happens on every single question, live, in the request path.
// INGESTION — runs whenever a help article is published or edited, not per-question
async function ingestArticle(article) {
const vector = await embed(article.title + "\n" + article.body);
await vectorStore.upsert({ id: article.id, vector, text: article.body, title: article.title });
}
// RETRIEVAL — runs on every question
async function retrieveContext(question) {
const questionVector = await embed(question);
return vectorStore.query(questionVector, { topK: 4 }); // 4 closest articles
}
The final piece stitches retrieval into the prompt, replacing the "paste the whole help center" version from two chapters back with only the handful of passages that actually matter to this question:
async function askLoopRAG(question) {
const matches = await retrieveContext(question);
const context = matches
.map((m, i) => `[${i + 1}] ${m.title}\n${m.text}`)
.join("\n\n");
const messages = [
{ role: "system", content: `${SYSTEM_PROMPT}\n\nUse only the numbered sources below to answer. Cite the source number(s) you actually used, like [1]. If none of the sources address the question, say so.\n\n${context}` },
{ role: "user", content: question },
];
return callModel(messages);
}
Ask Owen's question against this version and "Cancelling Your Subscription" comes back as the top match, its actual text — "cancelling stops future billing; there is no refund for the current billing period" — lands in the prompt, and the model answers correctly, citing [1]. That's real grounding, and it's a genuinely different mechanism from the system-prompt fix in chapter one: that fix taught the model to say "I don't know" when it had nothing; this teaches it to have something, retrieved fresh, specific to this exact question, at a fraction of the token cost of sending everything.
It is not, on its own, a guarantee of correctness — a model can still be handed the right passage and answer wrong anyway, which is exactly what the hallucination chapter is about. What it does guarantee is that the model now has access to the truth. Chapters seven and eight are both about what still goes wrong even with that access in hand.