mirror of
https://github.com/nlohmann/json.git
synced 2026-09-30 22:15:19 +00:00
get_codepoint() read the four hex digits of a \u escape with four calls to get(), each classified by a chain of range comparisons. For contiguous input, get_codepoint_bulk() now decodes them with one lookup per byte (hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a 256-entry table maps a byte to its value, or 0xFF for anything else, and an invalid digit shows in the OR of the four values. It then skips the four bytes and updates the position counters as four get() calls would. If a digit is invalid or fewer than four bytes are left, it changes nothing and the existing loop runs, so errors are reported with the same message and position as before. json::parse, best of 5 runs in separate processes (M1 Max): the escaped twitter.json (every non-ASCII character as \u) -13.6%, all other files within 0.3%. Tests compare the contiguous and the streaming path (value or exception message) for valid escapes, surrogate pairs, truncated and invalid digits at every position, and 3,000 seeded random escapes. Signed-off-by: Niels Lohmann <mail@nlohmann.me>