Skip to main content
CodeOath
← All posts

AI & LLM Engineering65 min total · 11 parts

AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling

Part 10 of 11 · ~3 min

Evaluating a Prompt Pipeline Like You'd Evaluate Code

Every fix in this reference so far was found by a real customer hitting a real gap, which is a fine way to discover a bug and a terrible way to run a production feature — Owen and Priti shouldn't be Ask Loop's test suite. A prompt is a function with an unusually wide input space and unusually variable output; treating it as untested-by-definition just because it's written in English rather than JavaScript is how all four bugs above make it to production in the first place.

The fix looks like ordinary testing, translated. Build a small, growing set of real questions with known-correct answers — a golden set — and run it automatically against every prompt or pipeline change, before it ships:

const GOLDEN_SET = [
  { question: "Can I get a refund after cancelling my annual plan?", expectContains: ["no refund", "stops future billing"], expectSource: "Cancelling Your Subscription" },
  { question: "What happens to extra projects when I downgrade from Team to Solo?", expectContains: ["archived", "90 days"], expectSource: "Downgrading Your Plan" },
  { question: "Do you offer a student discount?", expectEscalate: true },   // not in any doc — must say so, not guess
];

async function runGoldenSet() {
  const failures = [];
  for (const c of GOLDEN_SET) {
    const matches = await retrieveContext(c.question);
    if (c.expectSource && !matches.some((m) => m.title === c.expectSource)) {
      failures.push(`${c.question}: retrieval missed "${c.expectSource}"`);
      continue;
    }
    const reply = await askLoopGrounded(c.question);
    if (c.expectEscalate && !reply.escalate) failures.push(`${c.question}: should have escalated, didn't`);
    if (c.expectContains) {
      for (const phrase of c.expectContains) {
        if (!reply.content.toLowerCase().includes(phrase)) failures.push(`${c.question}: missing "${phrase}"`);
      }
    }
  }
  return failures;
}

Notice this checks two genuinely different things, and both matter independently. Retrieval quality — did the right document even get found — is checked separately from answer quality, because a wrong answer built on the right retrieved source is a prompt bug, and a wrong answer built on a missing or wrong source is a chunking or embedding bug, and confusing the two sends whoever's debugging it down the wrong path entirely.

The Priti downgrade bug from chapter five is exactly the kind of regression this catches automatically going forward: add her exact question to the golden set once, with the 90-day recovery window as a required phrase, and any future change to the chunker, the embedding model, or the system prompt that reintroduces the cut sentence fails a real, automated check the moment it happens — instead of waiting for the next real customer to hit it and ask, again, why their extra projects "disappeared."

An automated check like this one is necessarily shallow — string-containment isn't the same as genuine correctness, and it won't catch subtler drift in tone or a technically-accurate-but-misleading phrasing the way a person reading the transcript would. That's exactly why it's a floor, not the whole practice: a weekly human sample of real, anonymized conversations, read start to finish rather than grepped for keywords, is what catches the failures too subtle for a golden set to name in advance. Treat the two as complementary — the automated set as the fast, cheap gate that runs before every deploy, and periodic human review as the slower net underneath it.