Quick Answer: Among chunking strategies for RAG, recursive splitting on paragraph and sentence boundaries is the usual default; fixed-size cuts by length, semantic cuts where meaning shifts. A 2025 NAACL study found semantic chunking's extra computation wasn't justified by consistent gains. Switch to semantic or document-aware splits only when your own tests show they retrieve better than recursive ones.
Chunking is the step most RAG teams set once and forget. The splitter runs on framework defaults, the chatbot does fine in a demo, and months later nobody knows why it misses the answer in a table on page 40.
This guide compares the chunking strategies for RAG that engineers choose between: four options side by side, sourced starting sizes, a way to test them on your own questions, and the mistakes that quietly cost recall.
What is chunking in RAG?
Chunking in RAG is splitting source documents into smaller passages before they're embedded and indexed, so retrieval can return the few passages that answer a question instead of whole files. Every RAG chunking decision trades precision against context.
Embedding models accept limited input, and one vector for a 60-page manual blurs every topic into a single point. Small passages give retrieval something precise to match. A typical production pipeline:
- Parse each document and keep its structure: headings, lists, tables.
- Split it, usually recursively, into chunks of a few hundred tokens with modest overlap.
- Embed each chunk and store it in a vector index, often next to a keyword (BM25) index.
- At query time, retrieve the top matches, optionally re-rank them, and pass the best few to the model.
How do fixed-size, recursive and semantic chunking compare?
Fixed-size chunking is the cheapest and crudest, recursive chunking keeps paragraphs and sentences intact at almost no extra cost, and semantic chunking follows topic shifts at the price of extra embedding work. Document-aware chunking follows the file's own structure. Rule of thumb: recursive first, benchmark, then change.
| Strategy | How it splits | Setup effort | Indexing cost | Best for | Weak spot |
|---|---|---|---|---|---|
| Fixed-size | Every N tokens or characters | Lowest | One embedding per chunk | FAQs, chat logs, tickets | Cuts sentences and tables mid-way |
| Recursive | Paragraph breaks, then line breaks, then spaces, until chunks fit | Low; common framework default | One embedding per chunk | Most prose: docs, wikis, policies | Loses headings unless stored as metadata |
| Semantic | New chunk where similarity between neighboring sentences drops | Medium; needs a threshold | Extra embedding calls while splitting | Long text that changes topic often | Uneven sizes; gains not consistent in tests |
| Document-aware | Headings, sections, tables and code blocks | Medium to high; a parser per format | One embedding per chunk, plus parsing | Manuals, contracts, specs, Markdown | Only as good as each parser |
The evidence favors starting simple. A 2025 NAACL Findings paper tested semantic against fixed-size chunking on document retrieval, evidence retrieval and answer generation, and found semantic chunking's computational cost was not justified by consistent performance gains. That's a reason to measure first, not proof semantic splits never help. Choose among these RAG chunking strategies by the shape of your documents, not by how new the method is.
When does fixed-size chunking work well enough?
Fixed-size chunking works well enough when documents are short, uniform and self-contained: FAQ entries, support macros, product descriptions, chat transcripts. When each record holds one idea, a length-based split rarely cuts through meaning.
If you're setting up a chatbot knowledge base with minimal engineering effort, fixed or recursive chunking with default sizes is usually enough to start.
Chroma's technical report found a recursive splitter at 200 tokens with no overlap performed consistently well, though not best on every metric, and warned that some popular defaults performed relatively poorly. Whether to accept a hosted tool's defaults at all is covered in knowledge base tool vs custom RAG pipeline.
Default splits stop being enough when answers cite the right document but the wrong section, or tables come back cut in half.
When is semantic or document-aware chunking worth the extra work?
Semantic or document-aware chunking is worth the extra work on long, structured or mixed documents: annual reports, contracts, technical manuals, specifications, and anything where tables and headings carry meaning.
For these, document structure often matters more than semantic versus recursive. A contract already marks its clauses and a manual numbers its sections. Splitting on those boundaries and storing the section path as metadata usually beats inferring boundaries from embeddings.
LlamaIndex documents both routes: a semantic splitter that picks breakpoints by embedding similarity (by default, at the 95th percentile of dissimilarity), and a hierarchical parser that builds parent and child chunks, so its auto-merging retriever can match small chunks and swap in the larger parent when most of its children match. Semantic chunking earns its cost on long text with no reliable headings, such as transcripts or scraped pages.
What chunk size and overlap should you start with?
Start with chunks of a few hundred to about 1,000 tokens and an overlap of zero to 20 percent, then tune on your own questions. These are framework defaults and published results, not rules, and the unit matters: some splitters count characters, others count tokens.
| Source | Setting | Chunk size | Overlap | Unit |
|---|---|---|---|---|
| LlamaIndex API reference | SentenceSplitter default | 1,024 | 200 | tokens |
| LangChain API reference | RecursiveCharacterTextSplitter (inherits TextSplitter defaults) | 4,000 | 200 | characters |
| Chroma report, July 2024 | Strong all-round recursive result | 200 | 0 | tokens |
Framework defaults as documented on 25 September 2026.
LlamaIndex's SentenceSplitter defaults to 1,024-token chunks with a 200-token overlap, about 20 percent. LangChain calls its recursive character splitter the recommended one for generic text and counts characters, so 1,000 there is far smaller than 1,000 tokens in LlamaIndex.
As a starting point, go smaller (200 to 400 tokens) for precise facts such as part numbers or clauses, and larger (800 to 1,000) when answers need surrounding explanation.
How do you test chunking strategies for RAG on your own data?
Build a test set of real questions with the passages that should answer them, index the same corpus two or three ways, and compare retrieval before generated answers. The winner finds the right passage most often at an indexing cost and latency you can accept.
- Collect 50 to 100 real questions from tickets, search logs or subject experts, and mark each one's source passage.
- Index the corpus with two or three strategies, keeping the embedding model, top-k and re-ranker fixed.
- Measure recall@k (was a correct passage in the top k?) and precision@k (how much of what came back was relevant?).
- Score answer faithfulness on a sample: does each claim trace to a retrieved passage?
- Record indexing time, embedding calls, index size and query latency.
Change one setting at a time, and keep the test set: it lets you compare any strategy you consider later. Once you move on to grading generated answers, LLM judges versus human evaluators covers who should score them.
What mistakes should you avoid when chunking documents for RAG?
Whatever you choose among chunking strategies for RAG, the costly mistakes are one setting for every document type, stripping the context a chunk needs to be found, and never re-testing.
- One size for every document type. FAQs, contracts and API references need different sizes and splitters.
- Dropping headings and metadata. A chunk saying "this limit applies to all regions" can't match a question that names the product. Anthropic reported that prepending chunk-specific context before embedding and indexing cut top-20 retrieval failures by 49 percent, and 67 percent with re-ranking.
- Overlap that's too large. It inflates the index and fills the prompt with near-duplicates.
- Switching to semantic chunking on faith. Semantic isn't automatically better; measure first.
- Cutting tables apart. Keep tables whole or repeat the header in each piece.
- Never re-testing. New document types, a new embedding model or a new vector database for RAG all change which settings win.
How Origins AI chunks and indexes documents in Velocity AI Suite
Origins AI (originshq.com) handles this step in the Knowledge Foundation layer of Origins AI Velocity AI Suite, which its product page describes as data intake, document intelligence, knowledge structuring and retrieval indexing. The page reports support for 1,900+ data sources and 91+ document formats, lists Pinecone, Chroma, Weaviate and proprietary vector databases, and allows bring-your-own models.
The page doesn't publish a default chunk size or splitter, so settle chunking on your own documents: in a pilot, ask which settings each document format gets and see the retrieval test behind them.
For teams comparing AI knowledge base builders for chat and support, one ingested knowledge base feeds the suite's chat, voice and embedded delivery layer. The Origins AI products page says every product can run on-premise, in your own cloud account or air-gapped. With self-hosted models in on-premise and air-gapped modes, documents and queries stay inside your network; a hosted model receives the retrieved chunks with each query.
Talk to an engineer
Deciding how to chunk and index a real corpus? Bring sample documents and the questions your users ask, and an Origins AI engineer will walk through splitting, metadata and a retrieval test plan. Book a call.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


