OpenAI Plain-Text Token Counter
Paste raw text, code or Markdown to count its o200k_base tokens locally. This preset opens the encoding used for GPT-4o and GPT-4o mini in the OpenAI Cookbook. The complete string is tokenized with its original whitespace and Unicode composition. The result excludes message wrappers, tools, files, images and other API overhead, so it is not a count of an entire request.
Loading the selected local vocabulary and counting in a background worker…
- Plain-text tokens
- Counting…
Common questions
- Is this the official OpenAI token counter?
- It is a local utility using a pinned third-party BPE implementation of the named encoding. It is not an OpenAI account tool or an API call. Representative Latin, arithmetic and Japanese token IDs are checked against the official Cookbook fixtures. Use OpenAI’s request counter when you need all message and multimodal overhead.
- Can I use this for GPT-4o or GPT-4o mini text?
- The OpenAI Cookbook lists o200k_base for those models. This page counts their raw-text encoding, including spaces and punctuation, without creating a chat payload. The full request includes additional formatting and may include tools, schemas, files or images.
- Why does GPT-4 use a different count?
- GPT-4 and GPT-4 Turbo use cl100k_base in the Cookbook mapping. Choose that encoding or open the GPT-4 subpage. The vocabularies can split the same text differently; neither count is universally correct for every OpenAI model name.
- Does this tell me my API cost or ChatGPT usage limit?
- No. A raw-text count is not a price, a complete bill or an account usage report. This utility does not load prices, access your account or count cached, output, image, tool or reasoning tokens.
- Why show token IDs instead of colored character pieces?
- A token can end inside a UTF-8 character. Displaying each token as a decoded character can lose information, so this mode shows its actual IDs. Large previews show only the first 200 IDs, while the total includes every token in the paste.
Named BPE encodings count raw plain text locally. Complete API requests add message, tool and multimodal tokens that are not counted here. Claude, GPT-style and Llama heuristic estimates have no fixed measured error range and are not model-specific provider counts.