gizmobench
BrowseDaily

GPT-4 Plain-Text Token Counter

Count the raw text you intend to use with GPT-4 or GPT-4 Turbo using cl100k_base. This preset selects that local BPE encoding rather than o200k_base. Paste text, code or Markdown and inspect the total and token IDs. The counter preserves the original string, but it does not construct or count a complete chat request.

Plain-text token IDsCounting…

Loading the selected local vocabulary and counting in a background worker…

Plain-text tokens
Counting…

Raw plain text. The selected bundled BPE encoding counts the text value without further whitespace or Unicode normalization. Your browser may standardize pasted line endings in the text box. Special-token-looking strings are treated as ordinary text. This does not count complete API messages, roles, tools, schemas, PDFs, images, audio or account usage, and it does not determine a bill. Match the encoding to your model and use the provider's counter for a complete request.
Local text. Pasted text stays in this browser. The vocabularies are downloaded as site assets; counting does not send your text to an API. There is no fixed paste ceiling. Browser memory is the practical limit. Selected settings and drafts up to 50,000 characters can be remembered on this device.

Common questions

Which model names does this page’s encoding correspond to?
The OpenAI Cookbook maps GPT-4, GPT-4 Turbo and GPT-3.5 Turbo to cl100k_base, along with the listed text embedding models. This page counts raw text with that encoding. It does not resolve all model names or versions automatically.
Is GPT-4o the same as GPT-4 here?
No. The Cookbook maps GPT-4o and GPT-4o mini to o200k_base. Select that encoding for their raw text. Similar model names do not establish that the same vocabulary applies.
What happens to whitespace or special-token-looking strings?
The tokenizer keeps the text value’s whitespace and Unicode composition. The browser may standardize pasted line endings in the text box. A string such as a special token name is encoded as literal text, not treated as a control instruction. The result is a plain-text token count for the selected encoding.
Can I use the number as a complete API input count?
It does not include message roles, boundaries, system formatting, tools, schemas, files or images. The provider’s count for a complete request can differ. Use the request counter for a hard context limit or actual billing input.
Can I count a long prompt locally?
Yes. The complete paste is counted in a background worker with no fixed input ceiling. Browser memory is the practical limit. Only the token-ID preview is shortened; large drafts are omitted from local storage.

Named BPE encodings count raw plain text locally. Complete API requests add message, tool and multimodal tokens that are not counted here. Claude, GPT-style and Llama heuristic estimates have no fixed measured error range and are not model-specific provider counts.