Contact Us

Chunking Strategies for RAG Compared (2026)

Sep 25, 20267 min read
Blue light cut into equal segments feeding a glass lattice in a server room: Chunking Strategies for RAG Compared (2026)
chunking strategies for rag rag chunking rag chunking strategies

TL;DR

  • For long, structured documents such as contracts and manuals, splitting on headings and sections usually beats inferring boundaries from embeddings.
  • Start with chunks of a few hundred to about 1,000 tokens and an overlap of zero to 20 percent, then tune on your own questions.
  • Test two or three strategies on 50 to 100 real questions, keeping the embedding model fixed, and compare retrieval before generated answers.

Quick Answer: Among chunking strategies for RAG, recursive splitting on paragraph and sentence boundaries is the usual default; fixed-size cuts by length, semantic cuts where meaning shifts. A 2025 NAACL study found semantic chunking's extra computation wasn't justified by consistent gains. Switch to semantic or document-aware splits only when your own tests show they retrieve better than recursive ones.

Chunking is the step most RAG teams set once and forget. The splitter runs on framework defaults, the chatbot does fine in a demo, and months later nobody knows why it misses the answer in a table on page 40.

This guide compares the chunking strategies for RAG that engineers choose between: four options side by side, sourced starting sizes, a way to test them on your own questions, and the mistakes that quietly cost recall.

What is chunking in RAG?

Chunking in RAG is splitting source documents into smaller passages before they're embedded and indexed, so retrieval can return the few passages that answer a question instead of whole files. Every RAG chunking decision trades precision against context.

Embedding models accept limited input, and one vector for a 60-page manual blurs every topic into a single point. Small passages give retrieval something precise to match. A typical production pipeline:

  1. Parse each document and keep its structure: headings, lists, tables.
  2. Split it, usually recursively, into chunks of a few hundred tokens with modest overlap.
  3. Embed each chunk and store it in a vector index, often next to a keyword (BM25) index.
  4. At query time, retrieve the top matches, optionally re-rank them, and pass the best few to the model.

How do fixed-size, recursive and semantic chunking compare?

Fixed-size chunking is the cheapest and crudest, recursive chunking keeps paragraphs and sentences intact at almost no extra cost, and semantic chunking follows topic shifts at the price of extra embedding work. Document-aware chunking follows the file's own structure. Rule of thumb: recursive first, benchmark, then change.

Strategy How it splits Setup effort Indexing cost Best for Weak spot
Fixed-size Every N tokens or characters Lowest One embedding per chunk FAQs, chat logs, tickets Cuts sentences and tables mid-way
Recursive Paragraph breaks, then line breaks, then spaces, until chunks fit Low; common framework default One embedding per chunk Most prose: docs, wikis, policies Loses headings unless stored as metadata
Semantic New chunk where similarity between neighboring sentences drops Medium; needs a threshold Extra embedding calls while splitting Long text that changes topic often Uneven sizes; gains not consistent in tests
Document-aware Headings, sections, tables and code blocks Medium to high; a parser per format One embedding per chunk, plus parsing Manuals, contracts, specs, Markdown Only as good as each parser

The evidence favors starting simple. A 2025 NAACL Findings paper tested semantic against fixed-size chunking on document retrieval, evidence retrieval and answer generation, and found semantic chunking's computational cost was not justified by consistent performance gains. That's a reason to measure first, not proof semantic splits never help. Choose among these RAG chunking strategies by the shape of your documents, not by how new the method is.

When does fixed-size chunking work well enough?

Fixed-size chunking works well enough when documents are short, uniform and self-contained: FAQ entries, support macros, product descriptions, chat transcripts. When each record holds one idea, a length-based split rarely cuts through meaning.

If you're setting up a chatbot knowledge base with minimal engineering effort, fixed or recursive chunking with default sizes is usually enough to start.

Chroma's technical report found a recursive splitter at 200 tokens with no overlap performed consistently well, though not best on every metric, and warned that some popular defaults performed relatively poorly. Whether to accept a hosted tool's defaults at all is covered in knowledge base tool vs custom RAG pipeline.

Default splits stop being enough when answers cite the right document but the wrong section, or tables come back cut in half.

When is semantic or document-aware chunking worth the extra work?

Semantic or document-aware chunking is worth the extra work on long, structured or mixed documents: annual reports, contracts, technical manuals, specifications, and anything where tables and headings carry meaning.

For these, document structure often matters more than semantic versus recursive. A contract already marks its clauses and a manual numbers its sections. Splitting on those boundaries and storing the section path as metadata usually beats inferring boundaries from embeddings.

LlamaIndex documents both routes: a semantic splitter that picks breakpoints by embedding similarity (by default, at the 95th percentile of dissimilarity), and a hierarchical parser that builds parent and child chunks, so its auto-merging retriever can match small chunks and swap in the larger parent when most of its children match. Semantic chunking earns its cost on long text with no reliable headings, such as transcripts or scraped pages.

What chunk size and overlap should you start with?

Start with chunks of a few hundred to about 1,000 tokens and an overlap of zero to 20 percent, then tune on your own questions. These are framework defaults and published results, not rules, and the unit matters: some splitters count characters, others count tokens.

Source Setting Chunk size Overlap Unit
LlamaIndex API reference SentenceSplitter default 1,024 200 tokens
LangChain API reference RecursiveCharacterTextSplitter (inherits TextSplitter defaults) 4,000 200 characters
Chroma report, July 2024 Strong all-round recursive result 200 0 tokens

Framework defaults as documented on 25 September 2026.

LlamaIndex's SentenceSplitter defaults to 1,024-token chunks with a 200-token overlap, about 20 percent. LangChain calls its recursive character splitter the recommended one for generic text and counts characters, so 1,000 there is far smaller than 1,000 tokens in LlamaIndex.

As a starting point, go smaller (200 to 400 tokens) for precise facts such as part numbers or clauses, and larger (800 to 1,000) when answers need surrounding explanation.

How do you test chunking strategies for RAG on your own data?

Build a test set of real questions with the passages that should answer them, index the same corpus two or three ways, and compare retrieval before generated answers. The winner finds the right passage most often at an indexing cost and latency you can accept.

  1. Collect 50 to 100 real questions from tickets, search logs or subject experts, and mark each one's source passage.
  2. Index the corpus with two or three strategies, keeping the embedding model, top-k and re-ranker fixed.
  3. Measure recall@k (was a correct passage in the top k?) and precision@k (how much of what came back was relevant?).
  4. Score answer faithfulness on a sample: does each claim trace to a retrieved passage?
  5. Record indexing time, embedding calls, index size and query latency.

Change one setting at a time, and keep the test set: it lets you compare any strategy you consider later. Once you move on to grading generated answers, LLM judges versus human evaluators covers who should score them.

What mistakes should you avoid when chunking documents for RAG?

Whatever you choose among chunking strategies for RAG, the costly mistakes are one setting for every document type, stripping the context a chunk needs to be found, and never re-testing.

How Origins AI chunks and indexes documents in Velocity AI Suite

Origins AI (originshq.com) handles this step in the Knowledge Foundation layer of Origins AI Velocity AI Suite, which its product page describes as data intake, document intelligence, knowledge structuring and retrieval indexing. The page reports support for 1,900+ data sources and 91+ document formats, lists Pinecone, Chroma, Weaviate and proprietary vector databases, and allows bring-your-own models.

The page doesn't publish a default chunk size or splitter, so settle chunking on your own documents: in a pilot, ask which settings each document format gets and see the retrieval test behind them.

For teams comparing AI knowledge base builders for chat and support, one ingested knowledge base feeds the suite's chat, voice and embedded delivery layer. The Origins AI products page says every product can run on-premise, in your own cloud account or air-gapped. With self-hosted models in on-premise and air-gapped modes, documents and queries stay inside your network; a hosted model receives the retrieved chunks with each query.

Talk to an engineer

Deciding how to chunk and index a real corpus? Bring sample documents and the questions your users ask, and an Origins AI engineer will walk through splitting, metadata and a retrieval test plan. Book a call.

Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.

Frequently Asked Questions

What is the best chunk size for RAG?
There isn't one best size. A practical starting range is 200 to 1,000 tokens: LlamaIndex's sentence splitter defaults to 1,024, and Chroma's July 2024 report found a 200-token recursive split performed consistently well. Use the smaller end for fact lookups and the larger end for explanations, then confirm on your own questions.
What is semantic chunking?
Semantic chunking splits text where the meaning changes rather than at a fixed length. It embeds consecutive sentences, compares neighbors, and starts a new chunk where similarity drops past a threshold, producing topic-coherent chunks of uneven size. The trade-off is extra embedding work at indexing time, and published tests show the gain isn't consistent.
Do long-context models make chunking unnecessary?
Only for small collections. Anthropic's guidance is that a knowledge base under about 200,000 tokens, roughly 500 pages, can go straight into the prompt. Larger or fast-changing collections still need chunking, because sending everything on every query raises cost and latency, and retrieval points to the exact passage behind an answer.
How much chunk overlap should you use?
Start between zero and about 20 percent of the chunk size. LlamaIndex's default of 200 tokens on a 1,024-token chunk sits at the top of that range, while Chroma's best all-round recursive result used none. Add overlap when sentences run across boundaries, and cut it if near-duplicates fill your top results.
Does chunk size affect RAG cost?
Yes, in three places. Smaller chunks mean more embeddings to compute and store. Larger chunks mean more tokens sent to the model on every query, since the prompt carries the top-k chunks. Semantic chunking adds embedding calls while splitting, so track cost next to retrieval quality.
Should tables be chunked differently from text?
Yes. Separator-based splitters treat a table as ordinary text and can cut rows away from their header, leaving numbers with no labels. Keep small tables whole, split large ones into row groups that repeat the header, and store the caption and section heading as metadata.
Book a call

About the Author

Apoorva Kumar is Co-Founder and CEO of Origins AI (originshq.com), an AI engineering partner for product teams building AI workflows, AI agents and LLM integrations. A CSE graduate of IIT Kharagpur, Apoorva previously built and scaled technology at Sony, NuCash, YesMadam and FrontPage.