Skip to main content
CodeOath
← All posts

AI & LLM Engineering65 min total · 11 parts

AI & LLM Engineering Fundamentals: Prompting, RAG, Embeddings, and Function Calling

Part 6 of 11 · ~3 min

Chunking Strategy, and Where Retrieval Quietly Breaks

"Downgrading Your Plan" is one of Loopwork's longer articles, and here's the exact paragraph that matters:

You can downgrade at any time from Settings → Billing. Downgrading takes effect at the end of your current billing period. If you downgrade from Team to Solo, any projects beyond your new plan's project limit will be automatically archived, not deleted — you can restore them by upgrading again within 90 days.

A first pass at ingestion split every article into fixed 200-character chunks, with no regard for where a sentence happened to end, on the reasoning that shorter pieces retrieve more precisely than whole articles. Applied to that paragraph, it produced something close to this:

// Chunk A (retrieved): "...any projects beyond your new plan's project limit
// will be automatically archived, not"
//
// Chunk B (NOT retrieved — scored lower against this particular question):
// "deleted — you can restore them by upgrading again within 90 days."

A customer named Priti, mid-downgrade, asks Ask Loop what happens to her extra projects. Chunk A scores as the closer semantic match to her question and comes back as the only source. The answer she gets — "your extra projects beyond the new plan's limit will be automatically archived" — is true as far as it goes and drops the entire second half of the sentence: not deleted, and recoverable for 90 days. Nothing in the pipeline malfunctioned. Retrieval found a real match, and the model faithfully reported what was in it. The system failed at the boundary between two chunks, where a single sentence carrying the one qualifier that mattered most got cut in half and its two pieces went to different similarity scores.

This is the practical failure mode of fixed-size chunking: it treats text as an undifferentiated character stream and has no concept of "this clause depends on that one." Two changes fix it, and both are cheap relative to the cost of the bug above.

Chunk on structure, not character count. Split on paragraph or heading boundaries instead of an arbitrary character count, so a chunk boundary only ever falls somewhere a human would also consider a natural break — never mid-sentence, and ideally never mid-list either, since a numbered list with its intro sentence in one chunk and its actual items in the next is the same bug wearing a different shape.

function chunkByParagraph(articleBody) {
  return articleBody
    .split(/\n{2,}/)                 // blank-line-separated paragraphs
    .map((p) => p.trim())
    .filter((p) => p.length > 0);
}

Add overlap between adjacent chunks. Repeat the last sentence or two of one chunk at the start of the next, so a fact sitting right at a boundary shows up whole in at least one chunk even if the split still lands awkwardly:

function withOverlap(chunks, overlapSentences = 1) {
  return chunks.map((chunk, i) => {
    if (i === 0) return chunk;
    const prevTail = lastSentences(chunks[i - 1], overlapSentences);
    return `${prevTail} ${chunk}`;
  });
}

Reapply both to the downgrade paragraph and the archived-projects sentence stays intact end to end in a single chunk, because it's a complete unit under a structural split, and even if it hadn't been, the overlap would have carried the tail of the first chunk into the second. Priti's question now retrieves a chunk containing the full sentence, the model reports the 90-day recovery window along with the archiving, and nothing about the failure was visible anywhere except in a support transcript that read "technically true, meaningfully incomplete" — which is exactly the kind of bug that never shows up in a demo and only shows up once a real customer asks the one question a clean split happened to sever.