Write doubles with the shortest digits (Zmij)

dump() writes doubles with the conversion of Zmij by Victor Zverovich
(https://github.com/vitaut/zmij, MIT), ported to C++11 in
detail/conversions/zmij.hpp: the shortest decimal in the rounding
interval, the closest one if there are several. Grisu2 does not always
find the shortest digits; about 0.14% of random doubles are now written
differently (0.08% with fewer digits, 0.06% with the closest last
digit); short decimals such as 0.1 or 2555.56 are not affected. float
keeps Grisu2.

The layout of doubles is unchanged, but written differently: the digits
are converted eight at a time (the BCD conversion of Xiang JunBo, as in
Zmij) and stored with one byte swap per eight digits; leading and
trailing zeros are counted from those bytes; and the layouts of
format_buffer() are written with fixed-size moves instead of per-digit
loops and moves of the buffer (to_chars() uses a local buffer if the
caller's is shorter than the 41 bytes this may write).

The powers of ten come from the table for number parsing, adjusted
where it holds them rounded up, and from the compressed tables of Zmij
beyond 10^308. json::dump() gets faster on floats: canada -53%,
numbers -46%, mesh -37%, marine_ik -30%.

Tests: the powers of ten recomputed with a small big-integer; for random
doubles, all powers of two and of ten and their neighbors, and boundary
values: the output reads back as the same value, no decimal with one
digit fewer does, the layout equals that of format_buffer() for the same
digits, and (C++17) the digits equal those of std::to_chars.
The size ratios of canada.json in unit-binary_formats.cpp and one
expectation in unit-to_chars.cpp change with the shorter output.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-29 03:18:46 +02:00
parent cbdc502fbf
commit a04aa3d095
11 changed files with 1243 additions and 65 deletions

View File

@@ -33,7 +33,7 @@ TEST_CASE("Binary Formats" * doctest::skip())
const auto ubjson_2_size = json::to_ubjson(j, true).size();
const auto ubjson_3_size = json::to_ubjson(j, true, true).size();
CHECK(json_size == 2090303);
CHECK(json_size == 2090234);
CHECK(bjdata_1_size == 1112030);
CHECK(bjdata_2_size == 1224148);
CHECK(bjdata_3_size == 1224148);
@@ -46,16 +46,16 @@ TEST_CASE("Binary Formats" * doctest::skip())
CHECK(ubjson_3_size == 1169069);
CHECK((100.0 * double(json_size) / double(json_size)) == Approx(100.0));
CHECK((100.0 * double(bjdata_1_size) / double(json_size)) == Approx(53.199));
CHECK((100.0 * double(bjdata_2_size) / double(json_size)) == Approx(58.563));
CHECK((100.0 * double(bjdata_3_size) / double(json_size)) == Approx(58.563));
CHECK((100.0 * double(bon8_size) / double(json_size)) == Approx(50.509));
CHECK((100.0 * double(bson_size) / double(json_size)) == Approx(85.849));
CHECK((100.0 * double(cbor_size) / double(json_size)) == Approx(50.497));
CHECK((100.0 * double(msgpack_size) / double(json_size)) == Approx(50.526));
CHECK((100.0 * double(ubjson_1_size) / double(json_size)) == Approx(53.199));
CHECK((100.0 * double(ubjson_2_size) / double(json_size)) == Approx(58.563));
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(55.928));
CHECK((100.0 * double(bjdata_1_size) / double(json_size)) == Approx(53.201));
CHECK((100.0 * double(bjdata_2_size) / double(json_size)) == Approx(58.565));
CHECK((100.0 * double(bjdata_3_size) / double(json_size)) == Approx(58.565));
CHECK((100.0 * double(bon8_size) / double(json_size)) == Approx(50.511));
CHECK((100.0 * double(bson_size) / double(json_size)) == Approx(85.853));
CHECK((100.0 * double(cbor_size) / double(json_size)) == Approx(50.499));
CHECK((100.0 * double(msgpack_size) / double(json_size)) == Approx(50.528));
CHECK((100.0 * double(ubjson_1_size) / double(json_size)) == Approx(53.201));
CHECK((100.0 * double(ubjson_2_size) / double(json_size)) == Approx(58.565));
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(55.930));
}
SECTION("twitter.json")