mirror of
https://github.com/nlohmann/json.git
synced 2026-09-30 22:15:19 +00:00
An image is a document stored so that loading it needs no parsing: a 64-byte header, the nodes, the text, and the decoded strings (little-endian; version 1). - save() writes an edited document in its current state, in document order (floats that are not finite become null, as in dump()); the same document always gives the same bytes - load(pointer, size) and load(const vector&) borrow the image; load(vector&&) keeps it without a copy. The nodes are copied (aligned, and editable); the hash indexes of large objects are rebuilt. - image_check::full checks everything the parser guarantees (structure, bounds, UTF-8, strings of the source, number tokens and their values); bounds checks structure and bounds, so that reading and serializing stay safe; none trusts the image. A malformed image or a failed check throws the new parse_error.116; saving a discarded document (or images on a big-endian target) throws the new type_error.320; images of 4 GiB or more out_of_range.416. As images checked for bounds only can hold any bytes, the general float conversion now checks the token's grammar (and locates the point and the exponent itself), the exponent loop of the layout conversion takes digits as unsigned, and the serializer validates each non-ASCII sequence it decodes, throwing what basic_json::dump() throws for invalid UTF-8. Parsed and edited documents are not affected. The idea of images comes from zero-copy formats such as FlatBuffers and YaFF, the check from FlatBuffers' Verifier; no code is taken from them. Tests: round trips with every check (small documents, test files, large objects, edited documents with every kind of edit), ownership, all errors, one corruption per rejection branch of the check, and 12,000 seeded random corruptions, which must be rejected or read safely. The fuzzer json_view_image_fuzzer uses each input as an image and as a JSON text. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
118 lines
6.1 KiB
Markdown
118 lines
6.1 KiB
Markdown
# Fuzz testing
|
|
|
|
Each parser of the library (JSON, BJData, BON8, BSON, CBOR, MessagePack, and UBJSON) can be fuzz tested. Currently,
|
|
[libFuzzer](https://llvm.org/docs/LibFuzzer.html) and [afl++](https://github.com/AFLplusplus/AFLplusplus) are supported.
|
|
|
|
Additionally, `parse_json_view_fuzzer` (`tests/src/fuzzer-parse_json_view.cpp`) cross-checks `json_document`/`json_view`
|
|
(the zero-copy, read-only view declared in `json_view.hpp`) against `basic_json` on the same JSON text: it asserts that
|
|
`json_document::accept` agrees with `json::accept`, that an accepted input materializes to the same value `json::parse`
|
|
produces, and that a rejected input makes both parsers throw with an identical `what()`. It takes plain JSON text, so it
|
|
reuses the `corpus_json` corpus (or, for the `make fuzz_testing_json_view` target below, `tests/data/json_tests`) rather
|
|
than a format of its own.
|
|
|
|
`json_view_image_fuzzer` (`tests/src/fuzzer-json_view_image.cpp`) tests the images of `json_document` (`save()` and
|
|
`load()`). It uses each input twice: as an image, which `load()` must either reject with `parse_error.116` or read
|
|
safely (with `image_check::full`, the document must also serialize to the JSON it reads as), and as a JSON text, whose
|
|
image must load and serialize to the same text. A corpus of images can be made from JSON files with a small program
|
|
that calls `json_document::parse(text).save()`; plain JSON files work as well.
|
|
|
|
## Corpus creation
|
|
|
|
For most effective fuzzing, a [corpus](https://llvm.org/docs/LibFuzzer.html#corpus) should be provided. A corpus is a
|
|
directory with some simple input files that cover several features of the parser and is hence a good starting point
|
|
for mutations.
|
|
|
|
```shell
|
|
TEST_DATA_VERSION=3.2.0
|
|
wget https://github.com/nlohmann/json_test_data/archive/refs/tags/v$TEST_DATA_VERSION.zip
|
|
unzip v$TEST_DATA_VERSION.zip
|
|
rm v$TEST_DATA_VERSION.zip
|
|
for FORMAT in json bjdata bon8 bson cbor msgpack ubjson
|
|
do
|
|
rm -fr corpus_$FORMAT
|
|
mkdir corpus_$FORMAT
|
|
find json_test_data-$TEST_DATA_VERSION -size -5k -name "*.$FORMAT" -exec cp "{}" "corpus_$FORMAT" \;
|
|
done
|
|
rm -fr json_test_data-$TEST_DATA_VERSION
|
|
```
|
|
|
|
The generated corpus can be used with both libFuzzer and afl++. The remainder of this documentation assumes the corpus
|
|
directories have been created in the `tests` directory.
|
|
|
|
## libFuzzer
|
|
|
|
To use libFuzzer, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. In the `tests` directory, call
|
|
|
|
```shell
|
|
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
|
|
```
|
|
|
|
This creates a fuzz tester binary for each parser that supports these
|
|
[command line options](https://llvm.org/docs/LibFuzzer.html#options).
|
|
|
|
In case your default compiler is not a Clang compiler that includes libFuzzer (Clang 6.0 or later), you need to set the
|
|
`CXX` variable accordingly. Note the compiler provided by Xcode (AppleClang) does not contain libFuzzer. Please install
|
|
Clang via Homebrew calling `brew install llvm` and add `CXX=$(brew --prefix llvm)/bin/clang` to the `make` call:
|
|
|
|
```shell
|
|
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer" CXX=$(brew --prefix llvm)/bin/clang
|
|
```
|
|
|
|
Then pass the corpus directory as command-line argument (assuming it is located in `tests`):
|
|
|
|
```shell
|
|
./parse_cbor_fuzzer corpus_cbor
|
|
```
|
|
|
|
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is dumped into
|
|
a file starting with `crash-`.
|
|
|
|
## afl++
|
|
|
|
To use afl++, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. It will be replaced by a `libAFLDriver.a` to
|
|
re-use the same code written for libFuzzer with afl++. Furthermore, set `afl-clang-fast++` as compiler.
|
|
|
|
```shell
|
|
CXX=afl-clang-fast++ make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
|
|
```
|
|
|
|
Then the fuzzer is called like this in the `tests` directory:
|
|
|
|
```shell
|
|
afl-fuzz -i corpus_cbor -o out -- ./parse_cbor_fuzzer
|
|
```
|
|
|
|
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is written to the
|
|
directory `out`.
|
|
|
|
## OSS-Fuzz
|
|
|
|
The library is further fuzz-tested 24/7 by Google's [OSS-Fuzz project](https://github.com/google/oss-fuzz). It uses
|
|
the same `fuzzers` target as above and also relies on the `FUZZER_ENGINE` variable. See the used
|
|
[build script](https://github.com/google/oss-fuzz/blob/master/projects/json/build.sh) for more information.
|
|
|
|
In case the build at OSS-Fuzz fails, an issue will be created automatically.
|
|
|
|
### Handling OSS-Fuzz reports
|
|
|
|
OSS-Fuzz files the crashes it finds in its own [issue tracker](https://issues.oss-fuzz.com), not on GitHub. So that
|
|
each report can be traced to the change that fixed it, and each fix to the report it answers, fixes follow these
|
|
conventions:
|
|
|
|
- **Reference the OSS-Fuzz issue in the pull request**, next to any GitHub issue it closes, as `OSS-Fuzz: <id>` (for
|
|
example, `OSS-Fuzz: 563659413`), and in the commit message. The ID alone does not disclose the crash. If the report
|
|
was triaged into a GitHub issue, link the OSS-Fuzz issue there too.
|
|
- **Turn the reproducer into a unit test.** Download the testcase from the OSS-Fuzz report, reduce it if possible, and
|
|
add it as a regression test to the unit test of the affected format (e.g., `tests/src/unit-bjdata.cpp`), with a
|
|
comment naming the OSS-Fuzz issue. This way the input is checked by every CI run rather than only by OSS-Fuzz, and
|
|
it stays covered even if OSS-Fuzz later closes the report as not reproducible.
|
|
- **Keep the fuzzer drivers and the unit tests in sync.** The round-trip checks of the UBJSON and BJData drivers are
|
|
also run on a fixed corpus in the unit tests (see `tests/src/round_trip_corpus.hpp` and the "round-trip invariants"
|
|
test cases), so a regression shows up in CI first. When a driver's checks change, change the unit tests with them.
|
|
- **Record in the report whether the bug shipped.** OSS-Fuzz asks whether a crash was a short-lived regression or
|
|
affects a released version; answer it when the fix is merged, as it decides whether the fix needs a release note or
|
|
a security advisory (see the [security policy](../.github/SECURITY.md)).
|
|
|
|
After the fix is merged, OSS-Fuzz re-runs the reproducer on its next build and marks the report as verified and
|
|
closed. If it does not, the fix is incomplete.
|