Text tools
Counts update as you type, and the batch operations below cover case conversion, blank-line removal, line deduplication and sorting.
Text content
Live statistics
Batch operations
Counting text correctly, and cleaning it in bulk
Counting looks trivial until you have to agree on what counts. A character is one code point in most cases, but emoji built from combining sequences and flags occupy several, which is why some platforms count them differently. Excluding spaces removes every whitespace character, not just the space bar, so it is the honest figure for “how much did the author actually type”.
Words are counted as runs of letters and digits, which matches how English is written. That definition breaks down for Chinese, Japanese and Korean, where words are not separated by spaces — for CJK text the character count is the meaningful number, which is why it is broken out separately here.
Bytes are what storage and network transfer actually cost, and UTF-8 is variable length: ASCII characters take one byte, most accented Latin characters two, common CJK characters three, and emoji four. A 1,000-character Chinese document is roughly 3 KB, not 1 KB — worth knowing when a database column or an API limit is specified in bytes.
The batch operations cover the cleanup work that otherwise ends up in a scratch script. Trimming removes stray whitespace at the ends of the text or of every line; removing blank lines collapses the gaps left by copy-paste; deduplicating keeps the first occurrence of each line, which is the quickest way to clean an exported list; and sorting or reversing handles order. Punctuation normalisation converts full-width CJK punctuation to its ASCII equivalent, which matters when text is going into code, a URL or a CSV header.
Every operation rewrites the text in place and can be undone one step, so it is safe to experiment: apply an operation, look at the result, and undo if it was not what you wanted. Nothing is uploaded at any point — the counting and editing happen entirely in this page.
Frequently asked questions
How are words counted?
Words are counted as runs of Latin letters and digits separated by separators, which suits English prose. CJK text is counted by character instead, because it has no spaces between words.
Why is the byte count larger than the character count?
UTF-8 uses a variable number of bytes per character: ASCII takes one byte, accented Latin characters two, most CJK characters three, and emoji four. The byte count is what storage and network transfer actually cost.
Can I undo a batch operation?
Yes — “Undo last operation” rolls back the most recent batch change, so you can compare before and after. It keeps one step of history on this page.