What you are looking at
JSON-LP is a JSON lexer and parser written from scratch in pure Python: a character-level lexer, a recursive-descent parser, a typed error taxonomy with positions, and a command-line tool with pipeline exit codes. The office above is not a re-implementation of it in JavaScript. It fetches lexer.py and parser.py from the repository, unmodified, and runs them in CPython inside your browser through Pyodide. What you see is the real code working.
Every run is also a differential test. The same text goes to Python's own json.loads, held to the RFC-strict rule that a document must be an object or an array, and the two verdicts are compared. They have agreed on every case ever fed to them, which is the claim this page lets you attack.
How the office sorts
- The sack is your text. It is opened one character at a time.
- The sorting frame is the lexer. Each character, or run of characters, is a letter dropped into a pigeonhole by its token type: braces and brackets, quotes, strings, escapes, colons and commas, the literals
true,falseandnull, whitespace. The lexer has no number type: a number is filed as a string, and the pigeonhole says so. A letter that fits no hole is a lexer error, and the office stops there. - The belt carries the sorted letters, in order, to the reader.
- The reader is the parser. It is recursive descent: one function per grammar rule, each expecting particular letters next and complaining, with a position, when the wrong one arrives. A trailing comma is exactly that complaint.
- The inspectors are the comparator. One stamps the repository parser's verdict, one stamps
json.loads. When both stamps match, the seal is applied. When they differ, the letter is returned to sender, in red, because a disagreement is a bug in one of the two.
Try this
- Open the trailing comma sack. Both inspectors refuse it, for their own reasons, and the seal still goes on: agreement includes agreeing to refuse.
- Open not json. The reader refuses a bare word at the top level, and so does the RFC-strict reference. Drop the strictness and you would have a disagreement; that is why the reference is pinned.
- Type into the letter yourself. Every keystroke re-sorts the whole sack a third of a second later, with timings for the lexer and the parser separately.
- Drop a real
.jsonfile on the page, up to 2 MB. If you ever get the red stamp, please open an issue with the file.
The manifest: measured, not simulated
The corpus benchmark is bench.py in the repository and the conformance run is conformance.py; both print their numbers, and the README says which machine produced the ones above. The known weak spot is published with them: a single multi-megabyte string token scans quadratically, and the fix, chunked scanning, is named but not yet done.