mirror of
https://github.com/nlohmann/json.git
synced 2026-10-08 18:05:19 +00:00
* Check in the fuzzers that parsing without exceptions agrees Each fuzzer driver now also parses its input with allow_exceptions = false. That call must never throw a parse_error, must return a discarded value where parsing with exceptions fails, and must return the same value where it succeeds. Values are compared by their dump(), because NaN is not equal to itself. A plain !is_discarded() assertion, as suggested in #3642, would never fail: the drivers parse with exceptions, so a result can never be discarded. tests/fuzzing.md describes the checks and notes that OSS-Fuzz and CIFuzz already run LeakSanitizer, because their default address sanitizer includes it. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Use JSON_HAS_RANGE_VIEW_CONVERSION in the range view regression tests #5728 combined the JSON_HAS_RANGES and MinGW conditions into JSON_HAS_RANGE_VIEW_CONVERSION, but three test guards still spelled them out. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Test the serializer's buffers at their boundaries The dump() indent overflow survived full line coverage because the tests grew its buffer by only one step. This adds tests that land exactly on, and one past, the limits of the other two serializer buffers: - write_buffer (1024 bytes): strings of 1023, 1024 and 1025 bytes at the top level, and of 1022 and 1023 bytes inside an array, so that both guards in put_string() are hit at their boundary. Each is checked for dump() and for stream output. - string_buffer (512 bytes, flushed when fewer than 13 bytes remain): runs of two-byte escapes, and a surrogate pair written with 14 bytes of room, right after a flush, and one escape later. - The 8-byte bulk scan from the serializer side: 0 to 17 plain bytes followed by a quote, a control character, or a non-ASCII character. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Test the chunked string and binary reads of all binary formats The binary readers read strings and binary values in chunks of 4096 bytes. Only CBOR tested lengths around that size. MessagePack, UBJSON, BJData and BSON now round-trip lengths 0, 1, 4095, 4096, 4097, 8192 and 100000 from vector and pointer input, and must report a truncated payload as a parse error. UBJSON reads binary values as arrays of numbers, so it is tested with strings only. BJData binary values reach the chunked read only in Draft 3. BON8 decodes strings byte by byte and does not use this path. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
119 lines
5.9 KiB
Markdown
119 lines
5.9 KiB
Markdown
# Fuzz testing
|
|
|
|
Each parser of the library (JSON, BJData, BON8, BSON, CBOR, MessagePack, and UBJSON) can be fuzz tested. Currently,
|
|
[libFuzzer](https://llvm.org/docs/LibFuzzer.html) and [afl++](https://github.com/AFLplusplus/AFLplusplus) are supported.
|
|
|
|
## What the fuzzers check
|
|
|
|
Each fuzzer driver (`tests/src/fuzzer-parse_*.cpp`) parses its input twice: once with `allow_exceptions = false` and
|
|
once with exceptions. Both calls must agree. Where parsing with exceptions fails, the call without exceptions must
|
|
return a discarded value (or throw the same kind of non-parse error), and it must never throw a `parse_error`. Where
|
|
parsing succeeds, both calls must return the same value. The drivers then serialize the value, parse the result back,
|
|
and check that nothing was lost. The drivers check all of this with `assert`, so they refuse to build with `NDEBUG`.
|
|
|
|
## Corpus creation
|
|
|
|
For most effective fuzzing, a [corpus](https://llvm.org/docs/LibFuzzer.html#corpus) should be provided. A corpus is a
|
|
directory with some simple input files that cover several features of the parser and is hence a good starting point
|
|
for mutations.
|
|
|
|
```shell
|
|
TEST_DATA_VERSION=3.2.0
|
|
wget https://github.com/nlohmann/json_test_data/archive/refs/tags/v$TEST_DATA_VERSION.zip
|
|
unzip v$TEST_DATA_VERSION.zip
|
|
rm v$TEST_DATA_VERSION.zip
|
|
for FORMAT in json bjdata bon8 bson cbor msgpack ubjson
|
|
do
|
|
rm -fr corpus_$FORMAT
|
|
mkdir corpus_$FORMAT
|
|
find json_test_data-$TEST_DATA_VERSION -size -5k -name "*.$FORMAT" -exec cp "{}" "corpus_$FORMAT" \;
|
|
done
|
|
rm -fr json_test_data-$TEST_DATA_VERSION
|
|
```
|
|
|
|
The generated corpus can be used with both libFuzzer and afl++. The remainder of this documentation assumes the corpus
|
|
directories have been created in the `tests` directory.
|
|
|
|
## libFuzzer
|
|
|
|
To use libFuzzer, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. In the `tests` directory, call
|
|
|
|
```shell
|
|
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
|
|
```
|
|
|
|
This creates a fuzz tester binary for each parser that supports these
|
|
[command line options](https://llvm.org/docs/LibFuzzer.html#options).
|
|
|
|
In case your default compiler is not a Clang compiler that includes libFuzzer (Clang 6.0 or later), you need to set the
|
|
`CXX` variable accordingly. Note the compiler provided by Xcode (AppleClang) does not contain libFuzzer. Please install
|
|
Clang via Homebrew calling `brew install llvm` and add `CXX=$(brew --prefix llvm)/bin/clang` to the `make` call:
|
|
|
|
```shell
|
|
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer" CXX=$(brew --prefix llvm)/bin/clang
|
|
```
|
|
|
|
Then pass the corpus directory as command-line argument (assuming it is located in `tests`):
|
|
|
|
```shell
|
|
./parse_cbor_fuzzer corpus_cbor
|
|
```
|
|
|
|
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is dumped into
|
|
a file starting with `crash-`.
|
|
|
|
To also detect memory leaks, build with AddressSanitizer (`FUZZER_ENGINE="-fsanitize=fuzzer,address"`): libFuzzer then
|
|
runs LeakSanitizer by default (`-detect_leaks=1`). LeakSanitizer is not available with Apple Clang on macOS.
|
|
|
|
## afl++
|
|
|
|
To use afl++, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. It will be replaced by a `libAFLDriver.a` to
|
|
re-use the same code written for libFuzzer with afl++. Furthermore, set `afl-clang-fast++` as compiler.
|
|
|
|
```shell
|
|
CXX=afl-clang-fast++ make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
|
|
```
|
|
|
|
Then the fuzzer is called like this in the `tests` directory:
|
|
|
|
```shell
|
|
afl-fuzz -i corpus_cbor -o out -- ./parse_cbor_fuzzer
|
|
```
|
|
|
|
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is written to the
|
|
directory `out`.
|
|
|
|
## OSS-Fuzz
|
|
|
|
The library is further fuzz-tested 24/7 by Google's [OSS-Fuzz project](https://github.com/google/oss-fuzz). It uses
|
|
the same `fuzzers` target as above and also relies on the `FUZZER_ENGINE` variable. See the used
|
|
[build script](https://github.com/google/oss-fuzz/blob/master/projects/json/build.sh) for more information. Its default
|
|
`address` sanitizer includes LeakSanitizer, so OSS-Fuzz and the CIFuzz workflow (`.github/workflows/cifuzz.yml`) report
|
|
memory leaks, too.
|
|
|
|
In case the build at OSS-Fuzz fails, an issue will be created automatically.
|
|
|
|
### Handling OSS-Fuzz reports
|
|
|
|
OSS-Fuzz files the crashes it finds in its own [issue tracker](https://issues.oss-fuzz.com), not on GitHub. So that
|
|
each report can be traced to the change that fixed it, and each fix to the report it answers, fixes follow these
|
|
conventions:
|
|
|
|
- **Reference the OSS-Fuzz issue in the pull request**, next to any GitHub issue it closes, as `OSS-Fuzz: <id>` (for
|
|
example, `OSS-Fuzz: 563659413`), and in the commit message. The ID alone does not disclose the crash. If the report
|
|
was triaged into a GitHub issue, link the OSS-Fuzz issue there too.
|
|
- **Turn the reproducer into a unit test.** Download the testcase from the OSS-Fuzz report, reduce it if possible, and
|
|
add it as a regression test to the unit test of the affected format (e.g., `tests/src/unit-bjdata.cpp`), with a
|
|
comment naming the OSS-Fuzz issue. This way the input is checked by every CI run rather than only by OSS-Fuzz, and
|
|
it stays covered even if OSS-Fuzz later closes the report as not reproducible.
|
|
- **Keep the fuzzer drivers and the unit tests in sync.** The round-trip checks of the BJData, BON8, BSON, CBOR,
|
|
MessagePack and UBJSON drivers are also run on a fixed corpus in the unit tests (see
|
|
`tests/src/round_trip_corpus.hpp` and the "round-trip invariants" test cases), so a regression shows up in CI
|
|
first. When a driver's checks change, change the unit tests with them.
|
|
- **Record in the report whether the bug shipped.** OSS-Fuzz asks whether a crash was a short-lived regression or
|
|
affects a released version; answer it when the fix is merged, as it decides whether the fix needs a release note or
|
|
a security advisory (see the [security policy](../.github/SECURITY.md)).
|
|
|
|
After the fix is merged, OSS-Fuzz re-runs the reproducer on its next build and marks the report as verified and
|
|
closed. If it does not, the fix is incomplete.
|