What you are looking at
wc-Go is a rewrite of the Unix wc in Go that reads its input in chunks, counts lines, words, bytes and characters concurrently, and agrees with GNU coreutils on every case in a 59-case differential suite. The soroban above shows the four counts as beads. The real program reads 32 KB at a time; here the chunks are tiny, so you can watch the seams.
Each rod group is one counter. A soroban digit is read the way it has been read for four hundred years: the red bead above the beam is worth five when it is pushed down to the beam, and each bone bead below is worth one when pushed up. The digit is printed beneath the rods for anyone who would rather not read beads.
The seam problem
UTF-8 spells most non-ASCII characters in two, three or four bytes. Read a file in fixed-size chunks and sooner or later a chunk ends in the middle of one. Count the bytes naively and the character is either counted twice or, worse, decoded as garbage on both sides of the seam.
The algorithm looks back at most four bytes from the end of each chunk for the start of a character, decides whether the whole character arrived, and if it did not, holds those head bytes back and prepends them to the next chunk. Above, that is the blue bead moving from one tile to the next: the carry. The preset called boundary trap puts an é (bytes c3 a9) exactly across the 16-byte seam. Its first byte is carried; its second arrives with the next chunk; the character counter goes up once.
Word state has to survive the seam too. A word that starts in one chunk and continues in the next is one word, so the counter remembers whether it was inside a word when the chunk ended.
Try this
- Pick a chunk size of 4 on the frame and stream the emoji storm. Every emoji is four bytes, so almost every seam cuts one and the carry bead is rarely still.
- Stream utf-8 mix at 16. Devanagari, Cyrillic, Japanese and an emoji all count once each, and the bytes counter runs far ahead of the characters counter, which is the whole difference between
wc -candwc -m. - Drop a real file onto the page. Files over 4 KB stream at the real 32 KB chunk size, with throughput measured in this tab.
The tally: measured, not simulated
Agreement is checked by a differential suite that runs every flag combination through wc-Go and through coreutils and compares the output byte for byte. That suite caught a real bug during development: an invalid byte mid-stream made the counter silently drop the rest of a chunk, and the fix is the tail-splitting logic described above. It also caught the reference lying. Under a POSIX locale, coreutils wc -m counts bytes, not characters, so the suite pins LC_ALL=en_US.UTF-8 before it trusts a number.
To check it yourself: clone the repository, run go test ./... for the suite, and wc -lwmc yourfile beside wc-go yourfile should print the same four numbers.