mirror of
https://github.com/nlohmann/json.git
synced 2026-09-30 22:15:19 +00:00
Convert float and double with the library's own correctly rounded parser
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple arm64) are now converted by the library itself, correctly rounded and independent of the locale and of the C and C++ libraries: - The token is split into sign, significand w (at most 19 digits), and decimal exponent q, using the positions of the decimal point and the exponent that the scanners already recorded, so no character is classified again. - Clinger's fast path where w and 10^|q| are exact. - Eisel-Lemire otherwise, now templated for binary32 and binary64. - For tokens with more than 19 digits whose w and w + 1 round differently, an exact big-integer comparison with the midpoint between the two candidates (the digit comparison of fast_float, simplified). This replaces the separate token walks of Clinger's fast path and of Eisel-Lemire, the significant-digit gate that avoided the former, and, for float and double, std::from_chars and the locale-aware strtod. std::from_chars and strtold remain only for other long double formats (x87, binary128, double-double) and for types that are not IEEE-754. Values are bit-identical to before wherever the previous conversion was correctly rounded; tokens converted in a locale with a multi-byte decimal point are now also exact. Overflow still gives out_of_range.406, underflow a signed zero. convert_float() is the entry point for other parsers of JSON text: it converts like the lexer, without allocation for binary32/binary64. Tests: exact-bit tests for double and float (ties, subnormal and overflow boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for binary32, the round trips of 200,000 doubles and 100,000 floats without declines, 508 generated hard cases with the expected bits of both formats (float_hard_cases.hpp) through the converter and both scanners, and JSON-level overflow/underflow checks for double and float. The locale tests now check the values in a locale with a multi-byte decimal point. Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float credit now names the digit comparison. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
@@ -257,10 +257,11 @@ struct LocaleSwitchingSax final: public nlohmann::json_sax<json>
|
||||
|
||||
TEST_CASE("locale changes between lexer construction and number conversion (#5198)")
|
||||
{
|
||||
// The numbers are chosen so that the conversion also takes the strtod
|
||||
// fallback, which honors the locale that is current at conversion time:
|
||||
// too many significant digits for Clinger's fast path, an underflow that
|
||||
// std::from_chars rejects, and a plain value.
|
||||
// float and double are converted without the locale. A long double that
|
||||
// is not binary64 can take the strtold fallback, which honors the locale
|
||||
// that is current at conversion time. The numbers are chosen so that it
|
||||
// does: too many significant digits for Clinger's fast path, an underflow
|
||||
// that std::from_chars rejects, and a plain value.
|
||||
const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"};
|
||||
std::string text = "[";
|
||||
for (const auto& n : numbers)
|
||||
@@ -324,7 +325,8 @@ TEST_CASE("locale changes between lexer construction and number conversion (#519
|
||||
}
|
||||
}
|
||||
|
||||
// a long double goes through std::strtold unless std::from_chars supports it
|
||||
// a long double goes through std::strtold unless it is binary64 or
|
||||
// std::from_chars supports it
|
||||
{
|
||||
bool switched = false;
|
||||
const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept
|
||||
@@ -350,8 +352,9 @@ TEST_CASE("locale with a multi-byte decimal point")
|
||||
{
|
||||
// Some locales use a decimal point that is not a single character, e.g.
|
||||
// U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be
|
||||
// substituted in place for '.', so the strtod fallback stops early. The
|
||||
// conversion must still terminate rather than retry forever.
|
||||
// substituted in place for '.', so the strtold fallback (only for long
|
||||
// double formats other than binary64) stops early. The conversion must
|
||||
// still terminate rather than retry forever.
|
||||
const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}};
|
||||
bool tested = false;
|
||||
for (const char* name : names)
|
||||
@@ -369,12 +372,21 @@ TEST_CASE("locale with a multi-byte decimal point")
|
||||
tested = true;
|
||||
|
||||
// too many significant digits for Clinger's fast path, and an underflow
|
||||
// that std::from_chars rejects: both reach the strtod fallback
|
||||
// that std::from_chars rejects: double does not depend on the locale
|
||||
json j;
|
||||
CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]"));
|
||||
CHECK(j.is_array());
|
||||
CHECK(j[0] == 3.14159265358979323846);
|
||||
CHECK(j[1] == 0.0);
|
||||
CHECK(j[2] == -0.000123456789012345678);
|
||||
CHECK(json::accept("3.14159265358979323846"));
|
||||
|
||||
// a long double that reaches the strtold fallback must still terminate
|
||||
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
|
||||
long_double_json ld;
|
||||
CHECK_NOTHROW(ld = long_double_json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]"));
|
||||
CHECK(ld.is_array());
|
||||
|
||||
// a value the locale-independent paths convert is not affected
|
||||
CHECK(json::parse("12.5") == 12.5);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user