AI & LLM Engineering65 min total · 11 parts
AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling
Part 11 of 11 · ~3 min
Common Mistakes, and a Shipping Checklist
- No fallback for "I don't know." An instruction to "answer helpfully" implicitly rewards producing an answer over producing an honest one. Say explicitly, in the system prompt, that admitting insufficient information is a correct and expected outcome — not a failure to route around.
- Pasting entire documents into every prompt "to be safe." This is the forty-times cost spike from chapter two, and it also degrades answer quality — burying the one relevant paragraph inside 300,000 irrelevant words makes it measurably harder for the model to weight the right part correctly, on top of costing far more to do it.
- Fixed-size chunking with no overlap. The 90-day recovery clause that split across a chunk boundary in chapter five is not a rare edge case; any document with a clause, an exception, or a qualifier near where an arbitrary character count happens to land will produce the same bug. Split on structure, and overlap adjacent chunks.
- Trusting a model-supplied argument as authorization.
getInvoice(customerId, ...)withcustomerIdtaken from the model's own tool call, rather than from the authenticated session, is the exact shape of bug that lets one customer's conversation read another customer's data. The model's arguments get the same zero trust as any other unauthenticated client input. - Treating RAG as a hallucination cure. It fixes access to the right facts. It does not fix a model that fills a genuine gap in those facts with a plausible-sounding invention — that needs an explicit permission to say "not covered," and ideally a grounding check on the way out.
- Giving the model a tool that can take an irreversible or high-stakes action directly.
getPlanUsageandgetInvoiceare read-only and safe to hand over.issueRefundis not a tool at all here — it's a human-reviewed action the model can, at most, draft a request for. - Handing raw, undelimited user or document text straight into the prompt. Without an explicit "this is data, not instructions" boundary, a support ticket, an uploaded file, or an editable document is a place an attacker can hide a command the model has no built-in way to distinguish from a real one.
- Never testing prompts like code. A prompt change that isn't run against a golden set before shipping is a code change deployed with the test suite deleted — it happens to work in whatever three examples someone tried by hand and nothing stops the fourth from silently breaking.
None of the eight mechanisms above — prompting, tokens and context, embeddings, RAG, chunking, tool calling, hallucination handling, injection defense — is optional once real user data and real account actions are involved; skip any one of them and the corresponding bug in this reference is the one you'll meet in production, on a real support ticket, from a real customer. The authorization discipline in the function-calling and prompt-injection chapters is the same "never trust what the client claims about itself" principle covered in full in OAuth 2.0 and JWT Explained — worth reading if a tool call's authorization boundary in this piece felt underexplained. And building a golden-set harness that survives real refactors is exactly the discipline in Testing Fundamentals, applied to a system whose "code" happens to be written in English. For hands-on practice with the surrounding fundamentals — the async patterns, the API-calling code, the data structures a real retrieval pipeline leans on — the code lab is where to go build muscle memory rather than just read about it.
Continue learning
- Interview & Career PrepThe Non-Technical Half of the Interview: Behavioral Questions, the STAR Method, and What Recruiters Are Actually Scoring
- TypeScriptTypeScript Fundamentals: Types, Interfaces, Generics, and Why It Catches Bugs Before Runtime
- TestingTesting Fundamentals: Unit, Integration, and E2E Tests Done Right