Files
json/tests/benchmarks
Niels Lohmann 1e101ecac1 Add BON8 support (#2998)
* Add BON8 support

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments

- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Select the BON8 float prefix by type

get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename a test variable that Flawfinder mistakes for read()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures

- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BON8 strings in bulk from contiguous input

- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Link the BON8 functions from the other binary format pages

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the bulk scan flag after the input, not BON8

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BSON keys in bulk from contiguous input

BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures of the bulk-read tests

- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the explicit basic_json instantiation into its own test file

Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert the bytes of the BON8 test strings explicitly

The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 16:56:21 +02:00
..
2026-09-27 16:56:21 +02:00

Benchmarks

Micro-benchmarks for parsing, serialization and the binary formats, written with Google Benchmark. They are not run by CI; see When to run them.

What is measured

benchmark what it does
ParseFile, ParseString parse JSON from a file stream or a string
ParseIndented parse the large files re-indented by 4 spaces, for the lexer's whitespace handling
Dump serialize, compact (-) and indented (4)
ToCbor, BinaryToCbor write CBOR; BinaryToCbor writes binary values of growing size
FromMsgpack read MessagePack; unchanged over the years, so its numbers stay comparable across releases
FromBinaryBuffer, FromBinaryFile read CBOR, MessagePack, UBJSON, BJData and BSON from a buffer or a FILE*
FromBinaryShape read deeply nested, container-heavy and scalar-heavy documents in every binary format
FromCborChunkedString read CBOR strings split into indefinite-length chunks

The input files are those of nativejson-benchmark (canada, citm_catalog, twitter), a large jeopardy file, and number-heavy files (floats, signed_ints, ...). bytes_per_second counts the bytes read or written: the JSON text when parsing, the output when serializing.

Requirements

  • CMake 3.14 or later, a C++11 compiler, and Ninja for the make target.
  • Network access on the first configure: CMake downloads Google Benchmark and the test data into the build directory. To reuse a download of the test data, pass -DJSON_TestDataDirectory=<build directory>/test_files.
  • Google Benchmark is pinned to a release (1.9.5), so that results from different days stay comparable. To update it, change JSON_GOOGLE_BENCHMARK_VERSION and the archive's URL_HASH in CMakeLists.txt together.
  • The benchmarks include single_include/nlohmann/json.hpp, so run make amalgamate after changing anything in include/.

GCC and Clang builds use -O3 -flto -DNDEBUG.

Running them

From the repository root, this builds everything from scratch in cmake-build-benchmarks and runs all benchmarks:

make run_benchmarks

To build once and run selectively:

cmake -S tests/benchmarks -B build-benchmarks -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build-benchmarks
build-benchmarks/json_benchmarks --benchmark_filter='ParseString|Dump'

Useful options of json_benchmarks:

option effect
--benchmark_list_tests list the benchmarks instead of running them
--benchmark_filter=<regex> run only the benchmarks whose names match
--benchmark_repetitions=<n> run every benchmark n times and add mean, median, standard deviation and coefficient of variation
--benchmark_enable_random_interleaving=true run the repetitions in random order, which spreads out drifts such as thermal throttling
--benchmark_min_time=<seconds>s run each benchmark at least this long (e.g. 2s)
--benchmark_out=<file> --benchmark_out_format=json also write the results to a file, e.g. for compare.py

Reading the output

Each line shows the wall-clock Time and the CPU time per iteration, the number of Iterations Google Benchmark chose, and the throughput in bytes_per_second. With repetitions, the lines ending in _median are the ones to compare. A _cv (coefficient of variation) above a few percent means the machine was too noisy for small differences to mean anything.

Comparing two versions

To see what a change or a release did, build the same benchmarks twice: once against the header of the version to compare with, and once against the current one. JSON_BENCHMARK_INCLUDE_DIR names the directory holding the nlohmann/json.hpp to benchmark. For example, to compare the current checkout with 3.12.0:

# the header of the version to compare with
mkdir -p build-baseline-header/nlohmann
git show v3.12.0:single_include/nlohmann/json.hpp > build-baseline-header/nlohmann/json.hpp

# the same benchmarks, built against either header
cmake -S tests/benchmarks -B build-baseline -G Ninja -DCMAKE_BUILD_TYPE=Release \
      -DJSON_BENCHMARK_INCLUDE_DIR="$PWD/build-baseline-header"
cmake -S tests/benchmarks -B build-current -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build-baseline
cmake --build build-current

# run both, back to back
build-baseline/json_benchmarks --benchmark_repetitions=10 --benchmark_enable_random_interleaving=true \
                               --benchmark_out=build-baseline/results.json --benchmark_out_format=json
build-current/json_benchmarks --benchmark_repetitions=10 --benchmark_enable_random_interleaving=true \
                              --benchmark_out=build-current/results.json --benchmark_out_format=json

Google Benchmark ships a tool to compare the two result files. It needs NumPy and SciPy:

python3 -m venv build-venv
build-venv/bin/pip install numpy scipy
build-venv/bin/python build-current/_deps/benchmark-src/tools/compare.py -a benchmarks build-baseline/results.json build-current/results.json

The tool's own tools/requirements.txt pins NumPy and SciPy versions that need Python 3.11 or later; with an older Python, unpinned versions work as well. In its output:

  • the Time and CPU columns are relative changes: -0.35 means 35% faster, +0.10 means 10% slower;
  • _pvalue lines report a Mann-Whitney U test of whether the two versions differ. It needs at least 9 repetitions, and a p-value below 0.05 means the difference is unlikely to be noise;
  • OVERALL_GEOMEAN summarizes all benchmarks;
  • -a shows only the aggregates, not every repetition.

The header you compare with must support everything the benchmarks use. The current benchmarks build against 3.12.0. Only benchmarks present in both result files are compared, so for older releases, either filter the benchmarks or build that release's own tests/benchmarks against its own header.

Getting stable numbers

  • Build and run both versions on the same machine, one right after the other.
  • Keep the machine otherwise idle: no builds, no browser, and a laptop plugged in.
  • On Linux, set the CPU frequency governor to performance, e.g. sudo cpupower frequency-set --governor performance. Google Benchmark prints a warning when frequency scaling is enabled. Pinning the process to a core (taskset -c 2 ...) helps as well.
  • Use 10 or more repetitions with random interleaving, compare medians, and treat changes within the _cv as noise.

When to run them

They are a manual step, not part of CI: shared CI runners vary more between runs than most of the effects measured. Run the comparison above before a release, comparing the previous release tag with develop, and for pull requests that claim to change performance.