wc-Go

The shipped counting algorithm, worked on a soroban: bytes stream in, beads click up, and a character cut in half at a chunk seam is carried to the next chunk.
Counts agree with the reference.
Feed
Text
Stream log

What you are looking at

wc-Go is a rewrite of the Unix wc in Go that reads its input in chunks, counts lines, words, bytes and characters concurrently, and agrees with GNU coreutils on every case in a 59-case differential suite. The soroban above shows the four counts as beads. The real program reads 32 KB at a time; here the chunks are tiny, so you can watch the seams.

Each rod group is one counter. A soroban digit is read the way it has been read for four hundred years: the red bead above the beam is worth five when it is pushed down to the beam, and each bone bead below is worth one when pushed up. The digit is printed beneath the rods for anyone who would rather not read beads.

red bead, worth five at the beam
bone beads, worth one each, pushed up
the carry: head bytes of a split character
gold: the reference row agrees

The seam problem

UTF-8 spells most non-ASCII characters in two, three or four bytes. Read a file in fixed-size chunks and sooner or later a chunk ends in the middle of one. Count the bytes naively and the character is either counted twice or, worse, decoded as garbage on both sides of the seam.

The algorithm looks back at most four bytes from the end of each chunk for the start of a character, decides whether the whole character arrived, and if it did not, holds those head bytes back and prepends them to the next chunk. Above, that is the blue bead moving from one tile to the next: the carry. The preset called boundary trap puts an é (bytes c3 a9) exactly across the 16-byte seam. Its first byte is carried; its second arrives with the next chunk; the character counter goes up once.

Word state has to survive the seam too. A word that starts in one chunk and continues in the next is one word, so the counter remembers whether it was inside a word when the chunk ended.

Try this

The tally: measured, not simulated

Agreement is checked by a differential suite that runs every flag combination through wc-Go and through coreutils and compares the output byte for byte. That suite caught a real bug during development: an invalid byte mid-stream made the counter silently drop the rest of a chunk, and the fix is the tail-splitting logic described above. It also caught the reference lying. Under a POSIX locale, coreutils wc -m counts bytes, not characters, so the suite pins LC_ALL=en_US.UTF-8 before it trusts a number.

Differential agreement with coreutilsevery flag combination, byte for byte
59 of 59
Lines, on the repository's own harness3.9 to 12.8 times coreutils on specific operations
5.6 GB/s
Wordsthe one honest loss, published anyway
0.8 times

To check it yourself: clone the repository, run go test ./... for the suite, and wc -lwmc yourfile beside wc-go yourfile should print the same four numbers.