Skip to main content
CodeOath
← All posts

AI & LLM Engineering65 min total · 11 parts

AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling

Part 2 of 11 · ~5 min

Prompting Is an Engineering Discipline, Not a Trick

"Just ask it nicely" is the level Ask Loop is stuck at right now, and it's worth being precise about why that's not engineering: there's nothing to review, nothing to version, nothing that fails the same way twice, and no way to state what the function is actually supposed to guarantee. A prompt is not a clever sentence. It's the input to a function, and like any other input to a function, it deserves a spec.

The first real fix isn't a better question — it's a system prompt: a message with role: "system" that sits before anything the customer ever types, setting the ground rules for every turn that follows.

const SYSTEM_PROMPT = `You are Ask Loop, the support assistant embedded in Loopwork.

Rules:
- Only state a policy, price, deadline, or account fact if it was given to you
  directly in this conversation. Never rely on general knowledge of how other
  SaaS products typically handle something.
- If you don't have enough information to answer confidently, say so plainly
  and offer to connect the customer with a human. Do not guess.
- Keep answers under four sentences unless the customer asks for more detail.
- Never confirm that a refund, credit, or account change has been made. Only
  a human agent or an explicit tool result can confirm that.`;

async function askLoop(question) {
  return callModel([
    { role: "system", content: SYSTEM_PROMPT },
    { role: "user", content: question },
  ]);
}

Run Owen's refund question through this version and something better happens: the model says it doesn't have Loopwork's refund policy in front of it and offers to loop in a human. That's real progress, and it's worth noticing why it worked — not because the model suddenly "knows" the real policy (it still doesn't), but because the rule against guessing is now an explicit instruction sitting in the highest-priority position in the conversation, ahead of whatever the customer asks. It's not fixed, though. Ask Loop still can't state Loopwork's actual policy, because nobody has given it anything to state. That gap is the whole reason retrieval shows up two chapters from now — a system prompt can tell a model how to behave, but it cannot hand the model facts it was never given.

Roles, and what each one is actually for

Three roles do almost all the work in a normal conversation. system sets standing rules for the whole conversation and is written once, by you, never by the customer. user is whatever the human typed. assistant is what the model said last time — and this matters more than it looks like it should, because every earlier turn in a multi-turn conversation gets replayed back into the model as assistant messages on every subsequent call. The model has no memory between requests; what looks like "remembering" the last three exchanges is really you resending the entire transcript, system prompt and all, every single time.

async function askLoopTurn(history, question) {
  const messages = [
    { role: "system", content: SYSTEM_PROMPT },
    ...history,                              // every prior user/assistant pair, replayed
    { role: "user", content: question },
  ];
  const reply = await callModel(messages);
  return { reply, history: [...history, { role: "user", content: question }, reply] };
}

That "replay the whole transcript" detail isn't trivia — it's the reason the next chapter exists. A conversation that's gone on for forty turns is forty turns of tokens billed and re-processed on message forty-one, whether or not anything past turn five is still relevant.

Few-shot examples, for the cases a rule doesn't cover cleanly

A written rule ("don't guess") tells the model what to avoid. Showing it a worked example of the shape of a good refusal often does more than restating the rule more forcefully:

const FEW_SHOT = [
  { role: "user", content: "Do you offer student discounts?" },
  { role: "assistant", content: "I don't have information on a student discount in what I've been given — I don't want to guess and get that wrong. Want me to connect you with a human who can check?" },
];

Two or three of these, inserted right after the system prompt, tend to anchor tone and behavior harder than a longer paragraph of instructions — the model is a next-token predictor, and showing it what the next tokens after "I don't know" should look like is a more direct lever than describing them.

Temperature, and why Ask Loop should barely use it

temperature controls how much the model is willing to deviate from its single most-likely next token at each step. Near 0, it's close to deterministic — the same question tends to get the same answer. Turned up, it gets more varied, which is exactly what you want from something brainstorming marketing taglines and exactly what you don't want from something stating a company's cancellation policy. Ask Loop runs at 0.2 for a reason: a support assistant that gives a customer a different answer to the identical question twice in one afternoon has a trust problem before it has anything else.

None of this — not the system prompt, not the few-shot pairs, not the temperature setting — gives Ask Loop a single new fact. It's all shaping how the model uses what it already has. The next chapter is about the limits on how much you can hand it, and the one after that is about handing it the right facts at all.