Skip to main content
CodeOath
← All posts

AI & LLM Engineering65 min total · 11 parts

AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling

Part 4 of 11 · ~3 min

Embeddings and Semantic Search

Owen's actual question was "Can I get a refund if I cancel my annual plan?" The right help-center article is titled "Cancelling Your Subscription." Not one word overlaps between them. A keyword search for "refund" would miss that article entirely — it never uses the word "refund" anywhere, because Loopwork's policy is that cancellation simply stops future billing; there's no refund to describe. A keyword match would either return nothing, or worse, confidently return some unrelated article that happens to contain the literal word "refund" in a different context. What's needed is a match on meaning, and that's what an embedding gives you.

An embedding is a list of numbers — a vector — that a specialized model produces from a piece of text, positioned so that pieces of text with similar meaning end up as nearby vectors, and unrelated text ends up far apart. Real embeddings run to hundreds or thousands of dimensions; nobody can picture that directly, so here's the same idea in three dimensions, purely to see the mechanics:

// Toy 3-D stand-ins for real (1000+ dimension) embedding vectors — just
// to see the shape of the math. Each dimension here is standing in for
// some direction in meaning-space the real model learned on its own;
// nobody hand-picks what a real embedding's dimensions "mean."
const RESET_PASSWORD  = [0.90, 0.10, 0.05];
const CANCEL_ACCOUNT  = [0.05, 0.85, 0.10];
const EXPORT_TO_CSV   = [0.10, 0.05, 0.90];

function cosineSimilarity(a, b) {
  const dot = a.reduce((sum, x, i) => sum + x * b[i], 0);
  const magA = Math.sqrt(a.reduce((sum, x) => sum + x * x, 0));
  const magB = Math.sqrt(b.reduce((sum, x) => sum + x * x, 0));
  return dot / (magA * magB);
}

Cosine similarity measures the angle between two vectors, not their raw distance — two vectors pointing the same direction score close to 1 no matter how long either one is, which matters because it means document length doesn't skew the comparison the way raw distance would. Embed Owen's question — "Can I get a refund if I cancel my annual plan?" — and, because it's overwhelmingly about cancelling, it lands close to CANCEL_ACCOUNT in this toy space, not because the word "cancel" appears (it doesn't, verbatim, in "refund"), but because the model that produced these embeddings learned the relationship between the two ideas from enormous amounts of text where they show up together.

const questionVector = await embed("Can I get a refund if I cancel my annual plan?");
// questionVector ≈ [0.15, 0.82, 0.08] — close to CANCEL_ACCOUNT

cosineSimilarity(questionVector, CANCEL_ACCOUNT);   // ≈ 0.97 — best match
cosineSimilarity(questionVector, RESET_PASSWORD);   // ≈ 0.31
cosineSimilarity(questionVector, EXPORT_TO_CSV);    // ≈ 0.22

That's the entire trick underneath everything that follows: turn text into vectors with a consistent embedding model, and "which document is relevant to this question" becomes "which vector is closest," a problem with a clean mathematical answer instead of a fuzzy word-overlap heuristic. It's also worth being honest about the edges of it — semantic search finds what's conceptually near, not what's factually correct or complete, which is precisely the gap the hallucination chapter comes back to later. Closeness in meaning-space is not the same claim as being the right answer.