Skip to main content
CodeOath
← All posts

AI & LLM Engineering65 min total · 11 parts

AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling

Part 1 of 11 · ~2 min

Overview

Here is the version that shipped on a Friday afternoon, because a support lead asked for "something like ChatGPT, but for our help center" and someone had it working by lunch:

async function askLoop(question) {
  return callModel([{ role: "user", content: question }]);
}

Nine lines including the closing brace, and it genuinely works — type a question, get a fluent, confident, well-formatted answer back. Over the following Monday it told a customer named Owen that Loopwork offers a 30-day money-back guarantee on annual plans. Loopwork has never offered a refund on anything. The model wasn't being deceptive; it was doing exactly what it's built to do, which is produce the most statistically likely continuation of "Can I get a refund if I cancel my annual plan?" — and across everything it was trained on, the most likely continuation of that sentence is a description of a typical SaaS refund policy. It has no concept of this company's policy, because nobody ever told it, and it has no concept of the difference between "I know this" and "this is what usually comes next," because that distinction doesn't exist inside the model at all.

That gap — between fluent and correct — is the entire subject of this reference. Everything below gets built once, in order, onto the same feature: Ask Loop, a support assistant for Loopwork, a project-management SaaS with a growing help center, real customer accounts, and real billing data it should probably never touch carelessly. It starts exactly as broken as the nine lines above, and by the final chapter it retrieves grounded answers from the actual help docs, looks up a customer's real plan usage through a properly authorized tool call, admits when it doesn't know something instead of inventing an answer, resists a support ticket that tries to talk it into approving a refund it has no authority to approve, and has a test suite that catches a regression before a customer ever sees it.

One piece of vocabulary before any of that, because it's going to appear in nearly every code sample below:

// A thin wrapper around whichever provider Loopwork's backend calls.
// Every LLM API in wide use today accepts roughly this shape: an ordered
// list of messages, each tagged with who "said" it, and returns a
// generated continuation. The messages array IS the interface — nearly
// everything in this piece is really about what goes into that array.
async function callModel(messages, options = {}) {
  const res = await fetch(LLM_ENDPOINT, {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.LLM_API_KEY}` },
    body: JSON.stringify({ messages, temperature: 0.2, ...options }),
  });
  return (await res.json()).choices[0].message;
}

Keep that shape in mind — a list of role-tagged messages in, one generated message out. Every chapter from here is really a chapter about what belongs in that list, in what order, and where it comes from.