An image is a document stored so that loading it needs no parsing: a 64-byte header, the nodes, the text, and the decoded strings (little-endian; version 1). - save() writes an edited document in its current state, in document order (floats that are not finite become null, as in dump()); the same document always gives the same bytes - load(pointer, size) and load(const vector&) borrow the image; load(vector&&) keeps it without a copy. The nodes are copied (aligned, and editable); the hash indexes of large objects are rebuilt. - image_check::full checks everything the parser guarantees (structure, bounds, UTF-8, strings of the source, number tokens and their values); bounds checks structure and bounds, so that reading and serializing stay safe; none trusts the image. A malformed image or a failed check throws the new parse_error.116; saving a discarded document (or images on a big-endian target) throws the new type_error.320; images of 4 GiB or more out_of_range.416. As images checked for bounds only can hold any bytes, the general float conversion now checks the token's grammar (and locates the point and the exponent itself), the exponent loop of the layout conversion takes digits as unsigned, and the serializer validates each non-ASCII sequence it decodes, throwing what basic_json::dump() throws for invalid UTF-8. Parsed and edited documents are not affected. The idea of images comes from zero-copy formats such as FlatBuffers and YaFF, the check from FlatBuffers' Verifier; no code is taken from them. Tests: round trips with every check (small documents, test files, large objects, edited documents with every kind of edit), ownership, all errors, one corruption per rejection branch of the check, and 12,000 seeded random corruptions, which must be rejected or read safely. The fuzzer json_view_image_fuzzer uses each input as an image and as a JSON text. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
6.1 KiB
Fuzz testing
Each parser of the library (JSON, BJData, BON8, BSON, CBOR, MessagePack, and UBJSON) can be fuzz tested. Currently, libFuzzer and afl++ are supported.
Additionally, parse_json_view_fuzzer (tests/src/fuzzer-parse_json_view.cpp) cross-checks json_document/json_view
(the zero-copy, read-only view declared in json_view.hpp) against basic_json on the same JSON text: it asserts that
json_document::accept agrees with json::accept, that an accepted input materializes to the same value json::parse
produces, and that a rejected input makes both parsers throw with an identical what(). It takes plain JSON text, so it
reuses the corpus_json corpus (or, for the make fuzz_testing_json_view target below, tests/data/json_tests) rather
than a format of its own.
json_view_image_fuzzer (tests/src/fuzzer-json_view_image.cpp) tests the images of json_document (save() and
load()). It uses each input twice: as an image, which load() must either reject with parse_error.116 or read
safely (with image_check::full, the document must also serialize to the JSON it reads as), and as a JSON text, whose
image must load and serialize to the same text. A corpus of images can be made from JSON files with a small program
that calls json_document::parse(text).save(); plain JSON files work as well.
Corpus creation
For most effective fuzzing, a corpus should be provided. A corpus is a directory with some simple input files that cover several features of the parser and is hence a good starting point for mutations.
TEST_DATA_VERSION=3.2.0
wget https://github.com/nlohmann/json_test_data/archive/refs/tags/v$TEST_DATA_VERSION.zip
unzip v$TEST_DATA_VERSION.zip
rm v$TEST_DATA_VERSION.zip
for FORMAT in json bjdata bon8 bson cbor msgpack ubjson
do
rm -fr corpus_$FORMAT
mkdir corpus_$FORMAT
find json_test_data-$TEST_DATA_VERSION -size -5k -name "*.$FORMAT" -exec cp "{}" "corpus_$FORMAT" \;
done
rm -fr json_test_data-$TEST_DATA_VERSION
The generated corpus can be used with both libFuzzer and afl++. The remainder of this documentation assumes the corpus
directories have been created in the tests directory.
libFuzzer
To use libFuzzer, you need to pass -fsanitize=fuzzer as FUZZER_ENGINE. In the tests directory, call
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
This creates a fuzz tester binary for each parser that supports these command line options.
In case your default compiler is not a Clang compiler that includes libFuzzer (Clang 6.0 or later), you need to set the
CXX variable accordingly. Note the compiler provided by Xcode (AppleClang) does not contain libFuzzer. Please install
Clang via Homebrew calling brew install llvm and add CXX=$(brew --prefix llvm)/bin/clang to the make call:
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer" CXX=$(brew --prefix llvm)/bin/clang
Then pass the corpus directory as command-line argument (assuming it is located in tests):
./parse_cbor_fuzzer corpus_cbor
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is dumped into
a file starting with crash-.
afl++
To use afl++, you need to pass -fsanitize=fuzzer as FUZZER_ENGINE. It will be replaced by a libAFLDriver.a to
re-use the same code written for libFuzzer with afl++. Furthermore, set afl-clang-fast++ as compiler.
CXX=afl-clang-fast++ make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
Then the fuzzer is called like this in the tests directory:
afl-fuzz -i corpus_cbor -o out -- ./parse_cbor_fuzzer
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is written to the
directory out.
OSS-Fuzz
The library is further fuzz-tested 24/7 by Google's OSS-Fuzz project. It uses
the same fuzzers target as above and also relies on the FUZZER_ENGINE variable. See the used
build script for more information.
In case the build at OSS-Fuzz fails, an issue will be created automatically.
Handling OSS-Fuzz reports
OSS-Fuzz files the crashes it finds in its own issue tracker, not on GitHub. So that each report can be traced to the change that fixed it, and each fix to the report it answers, fixes follow these conventions:
- Reference the OSS-Fuzz issue in the pull request, next to any GitHub issue it closes, as
OSS-Fuzz: <id>(for example,OSS-Fuzz: 563659413), and in the commit message. The ID alone does not disclose the crash. If the report was triaged into a GitHub issue, link the OSS-Fuzz issue there too. - Turn the reproducer into a unit test. Download the testcase from the OSS-Fuzz report, reduce it if possible, and
add it as a regression test to the unit test of the affected format (e.g.,
tests/src/unit-bjdata.cpp), with a comment naming the OSS-Fuzz issue. This way the input is checked by every CI run rather than only by OSS-Fuzz, and it stays covered even if OSS-Fuzz later closes the report as not reproducible. - Keep the fuzzer drivers and the unit tests in sync. The round-trip checks of the UBJSON and BJData drivers are
also run on a fixed corpus in the unit tests (see
tests/src/round_trip_corpus.hppand the "round-trip invariants" test cases), so a regression shows up in CI first. When a driver's checks change, change the unit tests with them. - Record in the report whether the bug shipped. OSS-Fuzz asks whether a crash was a short-lived regression or affects a released version; answer it when the fix is merged, as it decides whether the fix needs a release note or a security advisory (see the security policy).
After the fix is merged, OSS-Fuzz re-runs the reproducer on its next build and marks the report as verified and closed. If it does not, the fix is incomplete.