gizmobench
BrowseDaily

Split Text by Word Count

Divide text into word-count chunks while retaining its original punctuation, spacing and line endings. This page opens in Words mode with 500 words per chunk and overlap set to zero. Change the size or overlap for your task, then select a numbered chunk to copy or save. Words here are separated by whitespace; this does not count model tokens or infer sentence meaning.

chunks19 files
  • The quick brown fox jumps over the lazy dog. Every chunk here ends on a character boundary, so n...

Chunks
19
Split by
Words per chunk
words
Overlap
words
Last chunk
64 words

19 chunks of up to 500 words, no overlap. The last one is 64 words.

Whitespace-separated words, with punctuation and original spacing retained. These are not model tokens. Copy and Save work on chunk-01.txt, the chunk selected in the list.

  • 48,120 chars at 4,000 eachno overlap
    13 chunks, the last is 120 characters
  • Boundary lands mid-charactera 25-byte emoji at a 40-byte budget
    26 bytes in the first chunk, not 40: the emoji stays whole
  • Overlap 200 on 4,000carry-over
    13 chunks, each repeating the previous 200 characters
What a chunk boundary can and cannot land on. Wherever this tool picks the boundary, by characters, by bytes, by lines or by words, it picks a grapheme cluster edge, the unit a reader calls a character, so the cut never falls inside an emoji, a skin tone, a flag, a keycap, a combining accent or a surrogate pair. In byte mode that means a chunk stops short of its budget rather than cutting a character: the sample this page opens with is 48,120 characters and 48,824 bytes as UTF-8, and the two numbers move apart exactly where the text stops being ASCII. A delimiter cut is yours to place, and it falls exactly where your delimiter does. Lines keep the ending they arrived with, and a delimiter stays at the end of the piece it closed, which is what lets the chunks add back up to the text they came from.

Up to 2,000,000 characters in one pass and up to 2,000 chunks, read and split here in your browser with nothing uploaded and no account. A delimiter may be typed with the usual escapes, so a blank line is \n\n and a tab is \t. The mode, the size, the overlap, the delimiter and the draft are kept in this browser alone, drafts up to 200,000 characters, and the Start over button above the tool forgets them.

Accuracy. Splitting is exact and reversible: join the chunks back in order with the overlap off and you get the original text back character for character, in every mode. Where this tool picks the boundary, by characters, by bytes or by lines or by words, the cut falls on a Unicode grapheme cluster edge (UAX#29), so it never lands inside an emoji, a combining mark or a surrogate pair, and a chunk stops short of a byte budget rather than breaking a character; a cut at a delimiter falls exactly where you put it, and an overlap must be smaller than the chunk or it is refused. The text is read and split in your browser.

Common questions

How do I split text by word count?
Choose Words, enter Words per chunk and set overlap to zero for separate chunks. A word is a run of non-whitespace graphemes; adjacent punctuation stays attached, and the original spaces and line endings are retained. Text without separating whitespace remains one word. This is not linguistic tokenisation or a model tokenizer.
Do the chunks really join back into the original?
Yes, in all five modes when overlap is zero. Lines retain their original endings, words retain spacing and punctuation, and delimiters stay at the end of the pieces they close. Concatenate the chunk text directly, without adding separators. File import validates UTF-8 and preserves a leading BOM; it refuses malformed byte sequences instead of replacing them.
How much text can it split at once?
Up to 2,000,000 grapheme characters and 2,000 chunks in one pass. Larger requests are refused without truncation. Files must contain valid UTF-8 text; convert another encoding in a text editor first. Each saved file is UTF-8 text. Plain splitting does not guarantee that chunks of CSV, JSON, Markdown or subtitles are independently valid documents.
Does it implement recursive or semantic splitting?
No. The modes count characters, bytes, words, existing lines or literal delimiter pieces. They do not choose paragraph or sentence boundaries by meaning, run embeddings, implement LangChain or LlamaIndex algorithms, or parse HTML and Markdown structure.
Does word overlap change reconstruction?
Yes. Overlap repeats whole word pieces, including their following spacing, in the next chunk. Use zero overlap when you want direct concatenation to restore the original.

Splitting is exact and reversible: join the chunks back in order with the overlap off and you get the original text back character for character, in every mode. Where this tool picks the boundary, by characters, by bytes or by lines or by words, the cut falls on a Unicode grapheme cluster edge (UAX#29), so it never lands inside an emoji, a combining mark or a surrogate pair, and a chunk stops short of a byte budget rather than breaking a character; a cut at a delimiter falls exactly where you put it, and an overlap must be smaller than the chunk or it is refused. The text is read and split in your browser.