mirror of
https://github.com/nlohmann/json.git
synced 2026-09-29 13:35:45 +00:00
* Add BON8 support Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format that uses the byte values that cannot begin a UTF-8 character as type markers, so strings need no length prefix. It is the most compact of the supported binary formats on the benchmark files. The reader is non-recursive like the other binary readers. A string ends at the first byte that cannot continue it, so the reader hands the one or two bytes it reads past a string back to the value that follows. The writer produces the canonical representation of the specification, except for NFC normalization; its output is identical to that of the reference implementation (HikoGUI) on all files of the test data. The round-trip tests need the .bon8 files of json_test_data 3.2.0. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Address review comments - Reuse detail::validate_one_utf8 to check strings in to_bon8; the error now names the first byte of the invalid sequence. - Document that to_bon8 leaves bytes in the output adapter on an exception, and that string_open is only an output of write_bon8_marker. - Explain why the pushback buffer of the BON8 reader cannot overflow. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Select the BON8 float prefix by type get_bon8_float_prefix only depends on the type of its argument, so make the type a template parameter instead of passing an unused value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Rename a test variable that Flawfinder mistakes for read() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures - compare the float in write_bon8_float with number_float_t constants, so GCC does not warn about a float-to-double conversion - mark check_bon8_utf8's context as used when exceptions are disabled - choose the compact float prefix in a helper rather than with nested conditional operators (clang-tidy) - use auto for the cast in the BON8 integer reader (clang-tidy) - write the int32 minimum test values as long long literals (MSVC C4146) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BON8 strings in bulk from contiguous input - copy the valid UTF-8 of a string in one step when the input is contiguous (twitter.json is read in 1.68 instead of 2.52 ms, jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack) - share the new valid_utf8_prefix() with the writer's UTF-8 check, which now skips ASCII 8 bytes at a time - let the fuzzer check that contiguous and stream input give the same value or error, and test both paths in the unit tests - clarify that a second 0xFF after a string is an empty string Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Link the BON8 functions from the other binary format pages Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the bulk scan flag after the input, not BON8 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BSON keys in bulk from contiguous input BSON keys (and array indices) are C-style strings, which were read byte by byte. For contiguous input they are now read up to their \x00-byte in one step, using the same bulk_scan flag as BON8 strings: twitter.json is read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of 3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys are almost all one-digit array indices, takes 2 % longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures of the bulk-read tests - skip the contiguous-versus-stream tests of BON8 strings and BSON keys when exceptions are disabled: they catch the parse errors of invalid input, and without exceptions the library aborts instead - use static_cast for the int64 test value (google-readability-casting) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the explicit basic_json instantiation into its own test file Linking test-regression3_cpp20 with clang and MinGW failed with "relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'", as test-regression2 did before #5511. The explicit instantiation of basic_json<> for #4825 compiles every member function, including the BON8 reader and writer, into that object, and it was already close to the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1, C++20). Give the instantiation a file of its own: unit-regression3 is now 1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built for the C++17 standard the regression was about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert the bytes of the BON8 test strings explicitly The str() helper constructed a std::string from a byte range, which converts each unsigned char implicitly; -fsanitize=integer reports that for bytes of 0x80 and above (ci_test_clang_sanitizer). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
105 lines
4.9 KiB
Markdown
105 lines
4.9 KiB
Markdown
# Fuzz testing
|
|
|
|
Each parser of the library (JSON, BJData, BON8, BSON, CBOR, MessagePack, and UBJSON) can be fuzz tested. Currently,
|
|
[libFuzzer](https://llvm.org/docs/LibFuzzer.html) and [afl++](https://github.com/AFLplusplus/AFLplusplus) are supported.
|
|
|
|
## Corpus creation
|
|
|
|
For most effective fuzzing, a [corpus](https://llvm.org/docs/LibFuzzer.html#corpus) should be provided. A corpus is a
|
|
directory with some simple input files that cover several features of the parser and is hence a good starting point
|
|
for mutations.
|
|
|
|
```shell
|
|
TEST_DATA_VERSION=3.2.0
|
|
wget https://github.com/nlohmann/json_test_data/archive/refs/tags/v$TEST_DATA_VERSION.zip
|
|
unzip v$TEST_DATA_VERSION.zip
|
|
rm v$TEST_DATA_VERSION.zip
|
|
for FORMAT in json bjdata bon8 bson cbor msgpack ubjson
|
|
do
|
|
rm -fr corpus_$FORMAT
|
|
mkdir corpus_$FORMAT
|
|
find json_test_data-$TEST_DATA_VERSION -size -5k -name "*.$FORMAT" -exec cp "{}" "corpus_$FORMAT" \;
|
|
done
|
|
rm -fr json_test_data-$TEST_DATA_VERSION
|
|
```
|
|
|
|
The generated corpus can be used with both libFuzzer and afl++. The remainder of this documentation assumes the corpus
|
|
directories have been created in the `tests` directory.
|
|
|
|
## libFuzzer
|
|
|
|
To use libFuzzer, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. In the `tests` directory, call
|
|
|
|
```shell
|
|
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
|
|
```
|
|
|
|
This creates a fuzz tester binary for each parser that supports these
|
|
[command line options](https://llvm.org/docs/LibFuzzer.html#options).
|
|
|
|
In case your default compiler is not a Clang compiler that includes libFuzzer (Clang 6.0 or later), you need to set the
|
|
`CXX` variable accordingly. Note the compiler provided by Xcode (AppleClang) does not contain libFuzzer. Please install
|
|
Clang via Homebrew calling `brew install llvm` and add `CXX=$(brew --prefix llvm)/bin/clang` to the `make` call:
|
|
|
|
```shell
|
|
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer" CXX=$(brew --prefix llvm)/bin/clang
|
|
```
|
|
|
|
Then pass the corpus directory as command-line argument (assuming it is located in `tests`):
|
|
|
|
```shell
|
|
./parse_cbor_fuzzer corpus_cbor
|
|
```
|
|
|
|
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is dumped into
|
|
a file starting with `crash-`.
|
|
|
|
## afl++
|
|
|
|
To use afl++, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. It will be replaced by a `libAFLDriver.a` to
|
|
re-use the same code written for libFuzzer with afl++. Furthermore, set `afl-clang-fast++` as compiler.
|
|
|
|
```shell
|
|
CXX=afl-clang-fast++ make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
|
|
```
|
|
|
|
Then the fuzzer is called like this in the `tests` directory:
|
|
|
|
```shell
|
|
afl-fuzz -i corpus_cbor -o out -- ./parse_cbor_fuzzer
|
|
```
|
|
|
|
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is written to the
|
|
directory `out`.
|
|
|
|
## OSS-Fuzz
|
|
|
|
The library is further fuzz-tested 24/7 by Google's [OSS-Fuzz project](https://github.com/google/oss-fuzz). It uses
|
|
the same `fuzzers` target as above and also relies on the `FUZZER_ENGINE` variable. See the used
|
|
[build script](https://github.com/google/oss-fuzz/blob/master/projects/json/build.sh) for more information.
|
|
|
|
In case the build at OSS-Fuzz fails, an issue will be created automatically.
|
|
|
|
### Handling OSS-Fuzz reports
|
|
|
|
OSS-Fuzz files the crashes it finds in its own [issue tracker](https://issues.oss-fuzz.com), not on GitHub. So that
|
|
each report can be traced to the change that fixed it, and each fix to the report it answers, fixes follow these
|
|
conventions:
|
|
|
|
- **Reference the OSS-Fuzz issue in the pull request**, next to any GitHub issue it closes, as `OSS-Fuzz: <id>` (for
|
|
example, `OSS-Fuzz: 563659413`), and in the commit message. The ID alone does not disclose the crash. If the report
|
|
was triaged into a GitHub issue, link the OSS-Fuzz issue there too.
|
|
- **Turn the reproducer into a unit test.** Download the testcase from the OSS-Fuzz report, reduce it if possible, and
|
|
add it as a regression test to the unit test of the affected format (e.g., `tests/src/unit-bjdata.cpp`), with a
|
|
comment naming the OSS-Fuzz issue. This way the input is checked by every CI run rather than only by OSS-Fuzz, and
|
|
it stays covered even if OSS-Fuzz later closes the report as not reproducible.
|
|
- **Keep the fuzzer drivers and the unit tests in sync.** The round-trip checks of the UBJSON and BJData drivers are
|
|
also run on a fixed corpus in the unit tests (see `tests/src/round_trip_corpus.hpp` and the "round-trip invariants"
|
|
test cases), so a regression shows up in CI first. When a driver's checks change, change the unit tests with them.
|
|
- **Record in the report whether the bug shipped.** OSS-Fuzz asks whether a crash was a short-lived regression or
|
|
affects a released version; answer it when the fix is merged, as it decides whether the fix needs a release note or
|
|
a security advisory (see the [security policy](../.github/SECURITY.md)).
|
|
|
|
After the fix is merged, OSS-Fuzz re-runs the reproducer on its next build and marks the report as verified and
|
|
closed. If it does not, the fix is incomplete.
|