Files
json/tests/fuzzing.md
Niels Lohmann 5876a22712 Add images of json_documents: save() and load()
An image is a document stored so that loading it needs no parsing: a
64-byte header, the nodes, the text, and the decoded strings
(little-endian; version 1).

- save() writes an edited document in its current state, in document
  order (floats that are not finite become null, as in dump()); the
  same document always gives the same bytes
- load(pointer, size) and load(const vector&) borrow the image;
  load(vector&&) keeps it without a copy. The nodes are copied (aligned,
  and editable); the hash indexes of large objects are rebuilt.
- image_check::full checks everything the parser guarantees (structure,
  bounds, UTF-8, strings of the source, number tokens and their values);
  bounds checks structure and bounds, so that reading and serializing
  stay safe; none trusts the image.

A malformed image or a failed check throws the new parse_error.116;
saving a discarded document (or images on a big-endian target) throws
the new type_error.320; images of 4 GiB or more out_of_range.416.

As images checked for bounds only can hold any bytes, the general float
conversion now checks the token's grammar (and locates the point and
the exponent itself), the exponent loop of the layout conversion takes
digits as unsigned, and the serializer validates each non-ASCII sequence
it decodes, throwing what basic_json::dump() throws for invalid UTF-8.
Parsed and edited documents are not affected.

The idea of images comes from zero-copy formats such as FlatBuffers and
YaFF, the check from FlatBuffers' Verifier; no code is taken from them.

Tests: round trips with every check (small documents, test files, large
objects, edited documents with every kind of edit), ownership, all
errors, one corruption per rejection branch of the check, and 12,000
seeded random corruptions, which must be rejected or read safely. The
fuzzer json_view_image_fuzzer uses each input as an image and as a JSON
text.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:20 +02:00

6.1 KiB

Fuzz testing

Each parser of the library (JSON, BJData, BON8, BSON, CBOR, MessagePack, and UBJSON) can be fuzz tested. Currently, libFuzzer and afl++ are supported.

Additionally, parse_json_view_fuzzer (tests/src/fuzzer-parse_json_view.cpp) cross-checks json_document/json_view (the zero-copy, read-only view declared in json_view.hpp) against basic_json on the same JSON text: it asserts that json_document::accept agrees with json::accept, that an accepted input materializes to the same value json::parse produces, and that a rejected input makes both parsers throw with an identical what(). It takes plain JSON text, so it reuses the corpus_json corpus (or, for the make fuzz_testing_json_view target below, tests/data/json_tests) rather than a format of its own.

json_view_image_fuzzer (tests/src/fuzzer-json_view_image.cpp) tests the images of json_document (save() and load()). It uses each input twice: as an image, which load() must either reject with parse_error.116 or read safely (with image_check::full, the document must also serialize to the JSON it reads as), and as a JSON text, whose image must load and serialize to the same text. A corpus of images can be made from JSON files with a small program that calls json_document::parse(text).save(); plain JSON files work as well.

Corpus creation

For most effective fuzzing, a corpus should be provided. A corpus is a directory with some simple input files that cover several features of the parser and is hence a good starting point for mutations.

TEST_DATA_VERSION=3.2.0
wget https://github.com/nlohmann/json_test_data/archive/refs/tags/v$TEST_DATA_VERSION.zip
unzip v$TEST_DATA_VERSION.zip
rm v$TEST_DATA_VERSION.zip
for FORMAT in json bjdata bon8 bson cbor msgpack ubjson
do
  rm -fr corpus_$FORMAT
  mkdir corpus_$FORMAT
  find json_test_data-$TEST_DATA_VERSION -size -5k -name "*.$FORMAT" -exec cp "{}" "corpus_$FORMAT" \;
done
rm -fr json_test_data-$TEST_DATA_VERSION

The generated corpus can be used with both libFuzzer and afl++. The remainder of this documentation assumes the corpus directories have been created in the tests directory.

libFuzzer

To use libFuzzer, you need to pass -fsanitize=fuzzer as FUZZER_ENGINE. In the tests directory, call

make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"

This creates a fuzz tester binary for each parser that supports these command line options.

In case your default compiler is not a Clang compiler that includes libFuzzer (Clang 6.0 or later), you need to set the CXX variable accordingly. Note the compiler provided by Xcode (AppleClang) does not contain libFuzzer. Please install Clang via Homebrew calling brew install llvm and add CXX=$(brew --prefix llvm)/bin/clang to the make call:

make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer" CXX=$(brew --prefix llvm)/bin/clang

Then pass the corpus directory as command-line argument (assuming it is located in tests):

./parse_cbor_fuzzer corpus_cbor

The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is dumped into a file starting with crash-.

afl++

To use afl++, you need to pass -fsanitize=fuzzer as FUZZER_ENGINE. It will be replaced by a libAFLDriver.a to re-use the same code written for libFuzzer with afl++. Furthermore, set afl-clang-fast++ as compiler.

CXX=afl-clang-fast++ make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer" 

Then the fuzzer is called like this in the tests directory:

afl-fuzz -i corpus_cbor -o out  -- ./parse_cbor_fuzzer 

The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is written to the directory out.

OSS-Fuzz

The library is further fuzz-tested 24/7 by Google's OSS-Fuzz project. It uses the same fuzzers target as above and also relies on the FUZZER_ENGINE variable. See the used build script for more information.

In case the build at OSS-Fuzz fails, an issue will be created automatically.

Handling OSS-Fuzz reports

OSS-Fuzz files the crashes it finds in its own issue tracker, not on GitHub. So that each report can be traced to the change that fixed it, and each fix to the report it answers, fixes follow these conventions:

  • Reference the OSS-Fuzz issue in the pull request, next to any GitHub issue it closes, as OSS-Fuzz: <id> (for example, OSS-Fuzz: 563659413), and in the commit message. The ID alone does not disclose the crash. If the report was triaged into a GitHub issue, link the OSS-Fuzz issue there too.
  • Turn the reproducer into a unit test. Download the testcase from the OSS-Fuzz report, reduce it if possible, and add it as a regression test to the unit test of the affected format (e.g., tests/src/unit-bjdata.cpp), with a comment naming the OSS-Fuzz issue. This way the input is checked by every CI run rather than only by OSS-Fuzz, and it stays covered even if OSS-Fuzz later closes the report as not reproducible.
  • Keep the fuzzer drivers and the unit tests in sync. The round-trip checks of the UBJSON and BJData drivers are also run on a fixed corpus in the unit tests (see tests/src/round_trip_corpus.hpp and the "round-trip invariants" test cases), so a regression shows up in CI first. When a driver's checks change, change the unit tests with them.
  • Record in the report whether the bug shipped. OSS-Fuzz asks whether a crash was a short-lived regression or affects a released version; answer it when the fix is merged, as it decides whether the fix needs a release note or a security advisory (see the security policy).

After the fix is merged, OSS-Fuzz re-runs the reproducer on its next build and marks the report as verified and closed. If it does not, the fix is incomplete.