Skip to main content
CodeOath
← All posts

AI & LLM Engineering65 min total · 11 parts

AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling

Part 3 of 11 · ~3 min

Tokens, Context Windows, and the Cost of a Careless Prompt

A well-meaning engineer's first idea for "just tell it the real policy" was the most direct one imaginable: paste all 300 help-center articles into the system prompt, so the model always has everything.

const SYSTEM_PROMPT_V2 = `${BASE_RULES}\n\nHere is our entire help center:\n${ALL_HELP_ARTICLES.join("\n\n")}`;

Within a week, Loopwork's LLM bill was running about forty times what it had been, and a fair number of Ask Loop's responses had gotten noticeably slower. Both problems trace back to the same unit: the token. A model doesn't read characters or words — it reads tokens, roughly a word or a word-fragment each, and every provider charges by token count, both for what you send in and for what the model generates back. Every one of those 300 articles, on every single question, meant paying to re-process the entire help center just to answer "how do I reset my password."

The context window is the second constraint, and it's not really about money — it's a hard ceiling. Every model has a maximum number of tokens it can hold across the system prompt, the full conversation history, anything else stuffed into the request, and the reply it's about to generate. Go over it and the request doesn't slow down; it fails outright, or the provider silently truncates the oldest content to make it fit — which, in a support conversation, can mean the system prompt itself gets pushed out first and Ask Loop quietly loses its own ground rules mid-conversation.

Put a number on it to see why "paste everything" can't be the answer at any scale. Loopwork's help center runs to roughly 300,000 words. A rough, widely-used rule of thumb is about three-quarters of a word per token for ordinary English, so that's somewhere north of 400,000 tokens — comfortably past what most context windows will hold at all, and even for a model with a large enough window, that's 400,000 tokens billed and re-processed on every single question, including "how do I reset my password," whose actual answer lives in one paragraph of one article.

The honest fix isn't a bigger window. It's sending less — specifically, sending only the handful of paragraphs actually relevant to this question, chosen at request time instead of hardcoded. That's a search problem: given a question, find the right slice of 300,000 words. Keyword search is the obvious first reach, and it's also exactly where the next chapter starts, with why it fails.

The conversation-history version of the same problem

The same discipline applies to history, not just documents. A support thread that's run twenty turns doesn't need every turn resent verbatim forever — a common, unglamorous fix is summarizing everything older than the last few turns into a short recap message and replaying that instead of the full transcript, trading a little fidelity on old context for a context budget that doesn't grow without bound. Worth flagging now, revisited properly once retrieval is in the mix, because retrieved context and conversation history are both competing for the same fixed budget of tokens.