mirror of
https://github.com/nlohmann/json.git
synced 2026-09-30 14:05:18 +00:00
* Add BON8 support Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format that uses the byte values that cannot begin a UTF-8 character as type markers, so strings need no length prefix. It is the most compact of the supported binary formats on the benchmark files. The reader is non-recursive like the other binary readers. A string ends at the first byte that cannot continue it, so the reader hands the one or two bytes it reads past a string back to the value that follows. The writer produces the canonical representation of the specification, except for NFC normalization; its output is identical to that of the reference implementation (HikoGUI) on all files of the test data. The round-trip tests need the .bon8 files of json_test_data 3.2.0. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Address review comments - Reuse detail::validate_one_utf8 to check strings in to_bon8; the error now names the first byte of the invalid sequence. - Document that to_bon8 leaves bytes in the output adapter on an exception, and that string_open is only an output of write_bon8_marker. - Explain why the pushback buffer of the BON8 reader cannot overflow. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Select the BON8 float prefix by type get_bon8_float_prefix only depends on the type of its argument, so make the type a template parameter instead of passing an unused value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Rename a test variable that Flawfinder mistakes for read() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures - compare the float in write_bon8_float with number_float_t constants, so GCC does not warn about a float-to-double conversion - mark check_bon8_utf8's context as used when exceptions are disabled - choose the compact float prefix in a helper rather than with nested conditional operators (clang-tidy) - use auto for the cast in the BON8 integer reader (clang-tidy) - write the int32 minimum test values as long long literals (MSVC C4146) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BON8 strings in bulk from contiguous input - copy the valid UTF-8 of a string in one step when the input is contiguous (twitter.json is read in 1.68 instead of 2.52 ms, jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack) - share the new valid_utf8_prefix() with the writer's UTF-8 check, which now skips ASCII 8 bytes at a time - let the fuzzer check that contiguous and stream input give the same value or error, and test both paths in the unit tests - clarify that a second 0xFF after a string is an empty string Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Link the BON8 functions from the other binary format pages Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the bulk scan flag after the input, not BON8 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BSON keys in bulk from contiguous input BSON keys (and array indices) are C-style strings, which were read byte by byte. For contiguous input they are now read up to their \x00-byte in one step, using the same bulk_scan flag as BON8 strings: twitter.json is read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of 3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys are almost all one-digit array indices, takes 2 % longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures of the bulk-read tests - skip the contiguous-versus-stream tests of BON8 strings and BSON keys when exceptions are disabled: they catch the parse errors of invalid input, and without exceptions the library aborts instead - use static_cast for the int64 test value (google-readability-casting) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the explicit basic_json instantiation into its own test file Linking test-regression3_cpp20 with clang and MinGW failed with "relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'", as test-regression2 did before #5511. The explicit instantiation of basic_json<> for #4825 compiles every member function, including the BON8 reader and writer, into that object, and it was already close to the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1, C++20). Give the instantiation a file of its own: unit-regression3 is now 1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built for the C++17 standard the regression was about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert the bytes of the BON8 test strings explicitly The str() helper constructed a std::string from a byte range, which converts each unsigned char implicitly; -fsanitize=integer reports that for bytes of 0x80 and above (ci_test_clang_sanitizer). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
217 lines
7.1 KiB
C++
217 lines
7.1 KiB
C++
// __ _____ _____ _____
|
|
// __| | __| | | | JSON for Modern C++ (supporting code)
|
|
// | | |__ | | | | | | version 3.12.0
|
|
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
|
|
//
|
|
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
|
|
// SPDX-License-Identifier: MIT
|
|
|
|
#include "doctest_compatibility.h"
|
|
|
|
#include <nlohmann/json.hpp>
|
|
using nlohmann::json;
|
|
|
|
#include <cstdint>
|
|
#include <limits>
|
|
#include <string>
|
|
#include <vector>
|
|
|
|
namespace
|
|
{
|
|
|
|
// a spread of values exercising every writer path: scalars of each width, the
|
|
// float paths, strings, binary, and containers big enough to reallocate
|
|
std::vector<json> test_values()
|
|
{
|
|
json big_array = json::array();
|
|
for (int i = 0; i < 5000; ++i)
|
|
{
|
|
big_array.push_back(i);
|
|
}
|
|
|
|
json big_object = json::object();
|
|
for (int i = 0; i < 1000; ++i)
|
|
{
|
|
big_object[std::to_string(i)] = i;
|
|
}
|
|
|
|
return
|
|
{
|
|
json(nullptr), json(true), json(false),
|
|
json(0), json(-1), json(255), json(-129), json(65535), json(-32769),
|
|
json(4294967295U), json(-2147483649LL), json(18446744073709551615ULL),
|
|
json(0.0), json(-0.5), json(3.1415926535897932),
|
|
json(""), json("hello"), json(std::string(1000, 'x')),
|
|
json::binary({0x00, 0x01, 0x02}, 42),
|
|
json::array(), json::object(),
|
|
json::array({1, 2, 3}), json({{"a", 1}, {"b", nullptr}}),
|
|
json({{"nested", {{"deep", json::array({1, "two", 3.0, nullptr})}}}}),
|
|
big_array, big_object
|
|
};
|
|
}
|
|
|
|
// BON8 has no integers above the int64 range, so to_bon8() rejects them
|
|
bool bon8_representable(const json& j)
|
|
{
|
|
return !j.is_number_unsigned() || j.get<std::uint64_t>() <= static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)());
|
|
}
|
|
|
|
// values to_bson() accepts: the document must be an object
|
|
std::vector<json> bson_values()
|
|
{
|
|
json big_object = json::object();
|
|
for (int i = 0; i < 1000; ++i)
|
|
{
|
|
big_object[std::to_string(i)] = i;
|
|
}
|
|
|
|
return
|
|
{
|
|
json::object(),
|
|
json({{"a", 1}, {"b", nullptr}, {"c", true}, {"d", 2.5}, {"e", "text"}}),
|
|
json({{"arr", json::array({1, 2, 3})}, {"obj", {{"k", "v"}}}}),
|
|
big_object
|
|
};
|
|
}
|
|
|
|
} // namespace
|
|
|
|
// The vector-returning to_*(j) overloads write through the non-virtual
|
|
// output_vector_sink, while to_*(j, adapter) goes through output_adapter_sink.
|
|
// The two are separate code paths that must stay byte-for-byte identical; these
|
|
// checks fail if either overload is ever changed without the other.
|
|
TEST_CASE("binary writer output sinks")
|
|
{
|
|
SECTION("vector sink and adapter sink agree")
|
|
{
|
|
// note: no SUBCASE inside these loops - doctest keys subcases by
|
|
// name/file/line, so a subcase in a loop body would only ever run for
|
|
// the first iteration
|
|
for (const auto& j : test_values())
|
|
{
|
|
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace));
|
|
|
|
std::vector<std::uint8_t> cbor;
|
|
json::to_cbor(j, cbor);
|
|
CHECK(json::to_cbor(j) == cbor);
|
|
|
|
std::vector<std::uint8_t> msgpack;
|
|
json::to_msgpack(j, msgpack);
|
|
CHECK(json::to_msgpack(j) == msgpack);
|
|
|
|
if (bon8_representable(j))
|
|
{
|
|
std::vector<std::uint8_t> bon8;
|
|
json::to_bon8(j, bon8);
|
|
CHECK(json::to_bon8(j) == bon8);
|
|
}
|
|
|
|
for (const bool use_size :
|
|
{
|
|
false, true
|
|
})
|
|
{
|
|
for (const bool use_type :
|
|
{
|
|
false, true
|
|
})
|
|
{
|
|
if (use_type && !use_size)
|
|
{
|
|
continue; // not a supported combination
|
|
}
|
|
CAPTURE(use_size);
|
|
CAPTURE(use_type);
|
|
std::vector<std::uint8_t> ubjson;
|
|
json::to_ubjson(j, ubjson, use_size, use_type);
|
|
CHECK(json::to_ubjson(j, use_size, use_type) == ubjson);
|
|
}
|
|
}
|
|
|
|
for (const auto version :
|
|
{
|
|
json::bjdata_version_t::draft2, json::bjdata_version_t::draft3
|
|
})
|
|
{
|
|
std::vector<std::uint8_t> bjdata;
|
|
json::to_bjdata(j, bjdata, false, false, version);
|
|
CHECK(json::to_bjdata(j, false, false, version) == bjdata);
|
|
}
|
|
}
|
|
|
|
for (const auto& j : bson_values())
|
|
{
|
|
CAPTURE(j.dump());
|
|
std::vector<std::uint8_t> bson;
|
|
json::to_bson(j, bson);
|
|
CHECK(json::to_bson(j) == bson);
|
|
}
|
|
}
|
|
|
|
SECTION("the char adapter produces the same bytes")
|
|
{
|
|
for (const auto& j : test_values())
|
|
{
|
|
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace));
|
|
|
|
const std::vector<std::uint8_t> expected = json::to_cbor(j);
|
|
std::vector<char> as_char;
|
|
json::to_cbor(j, as_char);
|
|
|
|
REQUIRE(as_char.size() == expected.size());
|
|
std::vector<std::uint8_t> as_bytes;
|
|
as_bytes.reserve(as_char.size());
|
|
for (const char c : as_char)
|
|
{
|
|
as_bytes.push_back(static_cast<std::uint8_t>(c));
|
|
}
|
|
CHECK(as_bytes == expected);
|
|
}
|
|
}
|
|
}
|
|
|
|
// binary_reserve_hint() is documented as a *lower* bound on the serialized size,
|
|
// so that reserving it up front can never leave the returned vector holding
|
|
// capacity beyond what the value actually needs.
|
|
TEST_CASE("binary_reserve_hint never over-reserves")
|
|
{
|
|
for (const auto& j : test_values())
|
|
{
|
|
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace));
|
|
|
|
const std::size_t hint = nlohmann::detail::binary_reserve_hint(j);
|
|
|
|
CHECK(hint <= json::to_cbor(j).size());
|
|
CHECK(hint <= json::to_msgpack(j).size());
|
|
CHECK(hint <= json::to_ubjson(j).size());
|
|
CHECK(hint <= json::to_ubjson(j, true, true).size());
|
|
CHECK(hint <= json::to_bjdata(j).size());
|
|
if (bon8_representable(j))
|
|
{
|
|
CHECK(hint <= json::to_bon8(j).size());
|
|
}
|
|
}
|
|
|
|
for (const auto& j : bson_values())
|
|
{
|
|
CAPTURE(j.dump());
|
|
CHECK(nlohmann::detail::binary_reserve_hint(j) <= json::to_bson(j).size());
|
|
}
|
|
|
|
SECTION("scalars get no hint")
|
|
{
|
|
CHECK(nlohmann::detail::binary_reserve_hint(json(nullptr)) == 0);
|
|
CHECK(nlohmann::detail::binary_reserve_hint(json(42)) == 0);
|
|
CHECK(nlohmann::detail::binary_reserve_hint(json("a string")) == 0);
|
|
CHECK(nlohmann::detail::binary_reserve_hint(json::binary({0x01})) == 0);
|
|
}
|
|
|
|
SECTION("containers are hinted from their element count")
|
|
{
|
|
CHECK(nlohmann::detail::binary_reserve_hint(json::array()) == 1);
|
|
CHECK(nlohmann::detail::binary_reserve_hint(json::array({1, 2, 3})) == 4);
|
|
CHECK(nlohmann::detail::binary_reserve_hint(json::object()) == 1);
|
|
CHECK(nlohmann::detail::binary_reserve_hint(json({{"a", 1}, {"b", 2}})) == 5);
|
|
}
|
|
}
|