diff --git a/README.md b/README.md index 9fd3f71fc..a74755820 100644 --- a/README.md +++ b/README.md @@ -1395,7 +1395,7 @@ THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR I - The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2009 [Florian Loitsch](https://florian.loitsch.com/) - The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/). - The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0). -- The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright © 2021 The fast_float authors +- The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright © 2021 The fast_float authors REUSE Software diff --git a/docs/mkdocs/docs/api/basic_json/number_float_t.md b/docs/mkdocs/docs/api/basic_json/number_float_t.md index 83c7011c5..5a404f53e 100644 --- a/docs/mkdocs/docs/api/basic_json/number_float_t.md +++ b/docs/mkdocs/docs/api/basic_json/number_float_t.md @@ -23,9 +23,10 @@ type to use. ## Template parameters `NumberFloatType` -: the type to store floating-point numbers. Parsing and serialization are implemented in terms of - `#!cpp std::strtof`/`#!cpp std::strtod`/`#!cpp std::strtold` and `#!cpp std::snprintf`, so the type must be - `#!cpp float`, `#!cpp double`, or `#!cpp long double`. The +: the type to store floating-point numbers. The parser converts `#!cpp float`, `#!cpp double`, and a + `#!cpp long double` that is IEEE 754 binary64 itself and other `#!cpp long double` formats with + `#!cpp std::from_chars` or `#!cpp std::strtold`, and serialization falls back to `#!cpp std::snprintf`, so the + type must be `#!cpp float`, `#!cpp double`, or `#!cpp long double`. The [binary formats](../../features/binary_formats/index.md) additionally require `#!cpp float` or `#!cpp double`, because they have no encoding for `#!cpp long double`. See [Template Parameter Requirements](../../features/types/template_parameters.md#numberfloattype). diff --git a/docs/mkdocs/docs/features/types/number_handling.md b/docs/mkdocs/docs/features/types/number_handling.md index cf37b044a..1bbf68e6d 100644 --- a/docs/mkdocs/docs/features/types/number_handling.md +++ b/docs/mkdocs/docs/features/types/number_handling.md @@ -71,10 +71,10 @@ otherwise, it uses unsigned integer storage. - Numbers with a decimal digit or scientific notation are always stored as `#!c double`. - The number types can be changed, see [Template number types](#template-number-types). - - As of version 3.9.1, the conversion is realized by - [`std::strtoull`](https://en.cppreference.com/w/cpp/string/byte/strtoul), - [`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and - [`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof), respectively. + - The library converts integers and floating-point numbers itself, independent of the locale. Floating-point + numbers are correctly rounded (to nearest, ties to even). Only a `#!c long double` that is not IEEE 754 binary64 + (e.g., the 80-bit x87 format) is converted with `#!cpp std::from_chars` where available, or with + [`std::strtold`](https://en.cppreference.com/w/cpp/string/byte/strtof). !!! example "Examples" @@ -85,10 +85,10 @@ otherwise, it uses unsigned integer storage. ### Number limits - Any 64-bit signed or unsigned integer can be stored without loss of precision. -- Numbers exceeding the limits of `#!c double` (i.e., numbers that after conversion via -[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) are not satisfying +- Numbers exceeding the limits of `#!c double` (i.e., numbers whose rounded value is not satisfying [`std::isfinite`](https://en.cppreference.com/w/cpp/numeric/math/isfinite) such as `#!c 1E400`) will throw exception -[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing. +[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing. Numbers too +small for `#!c double` (such as `#!c 1E-400`) become zero, with the sign of the number. - Floating-point numbers are rounded to the next number representable as `double`. For instance `#!c 3.141592653589793238462643383279` is stored as [`0x400921fb54442d18`](https://float.exposed/0x400921fb54442d18). This is the same behavior as the code `#!c double x = 3.141592653589793238462643383279;`. diff --git a/docs/mkdocs/docs/features/types/template_parameters.md b/docs/mkdocs/docs/features/types/template_parameters.md index 8ed23e156..5da5808ba 100644 --- a/docs/mkdocs/docs/features/types/template_parameters.md +++ b/docs/mkdocs/docs/features/types/template_parameters.md @@ -26,8 +26,9 @@ Requirements are split into two groups: diagnosed with dedicated error messages, and violating most of them results in a compiler error somewhere inside the library. Four violations are not caught at compile time at all: - - A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers, - because the lexer hands the buffer to `#!cpp std::strtoull`/`#!cpp std::strtoll`/`#!cpp std::strtod`. + - A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers + stored as a `#!cpp long double` that is not IEEE 754 binary64 (e.g., the 80-bit x87 format), because the lexer + hands the buffer to `#!cpp std::strtold`. - A stateful [`AllocatorType`](#allocatortype) compiles and silently ignores its state: allocation, deallocation, and [`get_allocator()`](../../api/basic_json/get_allocator.md) each use a different default-constructed instance. - The two [cross-specialization conversions](#cross-specialization-conversions) below. These abort on an assertion @@ -534,8 +535,10 @@ therefore silently changes parse results rather than raising an error. See `NumberFloatType` must be one of `#!cpp float`, `#!cpp double`, or `#!cpp long double`: -- The [parser](../parsing/index.md) converts number literals with `#!cpp std::strtof`, `#!cpp std::strtod`, or - `#!cpp std::strtold`; the library provides overloads for exactly these three types. +- The [parser](../parsing/index.md) converts number literals to `#!cpp float`, `#!cpp double`, and a + `#!cpp long double` that is IEEE 754 binary64 itself; other `#!cpp long double` formats are converted with + `#!cpp std::from_chars` where available, or with `#!cpp std::strtold`. The library provides overloads for exactly + these three types. - [`dump`](../../api/basic_json/dump.md) falls back to `#!cpp std::snprintf` with the `%g` and `%Lg` conversion specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads (`#!cpp float` is promoted to `#!cpp double`). diff --git a/docs/mkdocs/docs/home/license.md b/docs/mkdocs/docs/home/license.md index 863b7c1d6..7340c0fe2 100644 --- a/docs/mkdocs/docs/home/license.md +++ b/docs/mkdocs/docs/home/license.md @@ -20,4 +20,4 @@ The class contains a slightly modified version of the Grisu2 algorithm from Flor The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/). -The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright © 2021 The fast_float authors +The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright © 2021 The fast_float authors diff --git a/include/nlohmann/detail/input/lexer.hpp b/include/nlohmann/detail/input/lexer.hpp index 47de76c22..b072f3140 100644 --- a/include/nlohmann/detail/input/lexer.hpp +++ b/include/nlohmann/detail/input/lexer.hpp @@ -1059,9 +1059,11 @@ class lexer : public lexer_base token_type::parse_error otherwise @note The scanner is independent of the current locale: token_buffer - always holds `.`. Only the std::strtod fallback of convert_number() - depends on the locale, and it looks up the decimal point right - before converting (see detail::convert_float_locale_aware()). + always holds `.`. The conversion of float and double does not use + the locale either. Only the std::strtold fallback of + convert_number() for long double formats other than binary64 + depends on it, and it looks up the decimal point right before + converting (see detail::convert_float_locale_aware()). */ token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated. { @@ -1074,7 +1076,7 @@ class lexer : public lexer_base // offset just past the last mantissa byte in token_buffer (i.e. the // index of 'e'/'E', or the whole token when there is no exponent). - // convert_number() uses it to count significant digits; npos means + // convert_number() uses it to split the token; npos means // "not seen an exponent yet" and is resolved at scan_number_done std::size_t mantissa_end = std::string::npos; @@ -1404,8 +1406,8 @@ scan_number_done: @param[in] mantissa_end offset just past the last mantissa byte in token_buffer (the index of 'e'/'E', or token_buffer.size() when there is no exponent); - used to skip Clinger's fast path when it cannot - possibly succeed - see detail::mantissa_fits_clinger() + with decimal_point_position, it locates the parts + of a float token without scanning it again */ token_type convert_number(token_type number_type, std::size_t mantissa_end) { @@ -1474,10 +1476,11 @@ scan_number_done: } // this code is reached if we parse a floating-point number or if an - // integer conversion above overflowed. Prefer std::from_chars - // (Eisel-Lemire, locale-independent, correctly rounded) when available; - // otherwise the exact Clinger fast path (double only); otherwise the - // locale-aware strtof/strtod/strtold. + // integer conversion above overflowed. float and double (and long + // double where it is binary64) are converted by the library itself, + // correctly rounded and independent of the locale; other long double + // formats use std::from_chars when available, otherwise the + // locale-aware strtold. if (convert_float_fast(num_begin, num_end, decimal_point_position, mantissa_end, value_float)) { return token_type::value_float; diff --git a/include/nlohmann/detail/input/number_parse.hpp b/include/nlohmann/detail/input/number_parse.hpp index b85e758cb..5be5b7655 100644 --- a/include/nlohmann/detail/input/number_parse.hpp +++ b/include/nlohmann/detail/input/number_parse.hpp @@ -18,6 +18,7 @@ #include // memcpy #include // numeric_limits #include // string +#include // conditional, integral_constant, true_type, false_type #include #include @@ -35,10 +36,13 @@ #endif // This file contains the value-conversion helpers used by the lexer to turn an -// already-validated number token into a value, without the locale/errno -// overhead of std::strtoull/std::strtod where possible. They are free functions -// so the lexer stays focused on scanning (see lexer::convert_number()) and so -// that other parsers of JSON text can convert tokens exactly like it does. +// already-validated number token into a value. Integers and binary32/binary64 +// floats (float, double, and long double where it is binary64) are converted +// by the library itself, without the locale/errno overhead of +// std::strtoull/std::strtod and correctly rounded; other long double formats +// use std::from_chars or std::strtold. They are free functions so the lexer +// stays focused on scanning (see lexer::convert_number()) and so that other +// parsers of JSON text can convert tokens exactly like it does. NLOHMANN_JSON_NAMESPACE_BEGIN namespace detail @@ -115,198 +119,133 @@ bool parse_integer_signed(const char* first, const char* last, NumberIntegerType } /*! -@brief exact fast path for parsing a `double` (Clinger's algorithm) - -For the common case - at most 19 significant digits, a decimal exponent in -[-22, 22], and a significand below 2^53 - the value equals significand * -10^exp computed in IEEE-754 double arithmetic, which is exact under -round-to-nearest because both operands are exactly representable. This is the -same fast path used by fast_float/simdjson; the general cases are left to -std::strtod. The parser only activates for number_float_t == double; float and -long double keep the std::strtof/std::strtold paths (see the templated overload -below). - -@param[in] first pointer to the first character of the number -@param[in] last pointer past the last character -@param[out] out the parsed value on success -@return true if the value was parsed exactly; false to fall back to strtod +@brief parameters of the IEEE-754 binary32 and binary64 formats for the float + conversion (after fast_float's binary_format) */ -inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept +template +struct ieee_binary_format; + +template<> +struct ieee_binary_format<24> // binary32 { -#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0 - // Clinger's fast path is only exact when double operations are evaluated in - // true double precision. On platforms that keep intermediates in extended - // precision (e.g. the x87 FPU on 32-bit x86, where FLT_EVAL_METHOD == 2) the - // single significand * 10^scale step is double-rounded and can be 1 ULP off, - // so decline and let the caller fall back to the correctly-rounded - // std::from_chars / std::strtod path. - static_cast(first); - static_cast(last); - static_cast(out); - return false; -#else - static const std::array powers_of_ten = + static constexpr int mantissa_bits() noexcept { - { - 1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11, - 1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22 - } - }; + return 23; + } + static constexpr int sign_bit() noexcept + { + return 31; + } + static constexpr int minimum_exponent() noexcept + { + return -127; + } + static constexpr int infinite_power() noexcept + { + return 0xFF; + } + // w * 10^q with w < 2^64 is below half the smallest subnormal number for + // q < smallest_power_of_ten() and at least infinity for q > largest_power_of_ten() + static constexpr int smallest_power_of_ten() noexcept + { + return -64; + } + static constexpr int largest_power_of_ten() noexcept + { + return 38; + } + // w * 10^q can only be exactly between two numbers for q in this range + static constexpr int min_exponent_round_to_even() noexcept + { + return -17; + } + static constexpr int max_exponent_round_to_even() noexcept + { + return 10; + } + // Clinger's fast path: w and 10^|q| are exact + static constexpr int max_exponent_fast_path() noexcept + { + return 10; + } + static constexpr std::uint64_t max_mantissa_fast_path() noexcept + { + return std::uint64_t{2} << 23u; + } + // a midpoint between two numbers has at most this many significant digits + static constexpr std::int64_t max_digits() noexcept + { + return 114; + } +}; - const char* p = first; - bool negative = false; - if (p != last && (*p == '-' || *p == '+')) - { - negative = (*p == '-'); - ++p; - } - - std::uint64_t significand = 0; - int num_digits = 0; - int fractional_digits = 0; - bool seen_dot = false; - bool any_digit = false; - for (; p != last; ++p) - { - const char c = *p; - if (c >= '0' && c <= '9') - { - any_digit = true; - if (JSON_HEDLEY_UNLIKELY(num_digits >= 19)) - { - return false; // significand may not fit into uint64_t - } - significand = (significand * 10u) + static_cast(c - '0'); - ++num_digits; - fractional_digits += static_cast(seen_dot); - } - else if (c == '.') - { - if (JSON_HEDLEY_UNLIKELY(seen_dot)) - { - return false; - } - seen_dot = true; - } - else if (c == 'e' || c == 'E') - { - ++p; - break; - } - else - { - return false; - } - } - if (JSON_HEDLEY_UNLIKELY(!any_digit)) - { - return false; - } - - int exponent = 0; - if (p != last) // an exponent part remains - { - bool exp_negative = false; - if (p != last && (*p == '-' || *p == '+')) - { - exp_negative = (*p == '-'); - ++p; - } - bool any_exp_digit = false; - for (; p != last; ++p) - { - if (JSON_HEDLEY_UNLIKELY(*p < '0' || *p > '9')) - { - return false; - } - exponent = (exponent * 10) + (*p - '0'); - any_exp_digit = true; - if (JSON_HEDLEY_UNLIKELY(exponent > 9999)) - { - return false; - } - } - if (JSON_HEDLEY_UNLIKELY(!any_exp_digit)) - { - return false; - } - if (exp_negative) - { - exponent = -exponent; - } - } - - const int scale = exponent - fractional_digits; - if (JSON_HEDLEY_UNLIKELY(significand >= (static_cast(1) << 53))) - { - return false; // significand not exactly representable as double - } - - auto result = static_cast(significand); - if (scale >= 0) - { - if (JSON_HEDLEY_UNLIKELY(scale > 22)) - { - return false; - } - result *= powers_of_ten[static_cast(scale)]; - } - else - { - if (JSON_HEDLEY_UNLIKELY(-scale > 22)) - { - return false; - } - result /= powers_of_ten[static_cast(-scale)]; - } - out = negative ? -result : result; - return true; -#endif -} - -/// fast float path is only exact for `double`; decline for float/long double -template -bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept +template<> +struct ieee_binary_format<53> // binary64 { - return false; -} + static constexpr int mantissa_bits() noexcept + { + return 52; + } + static constexpr int sign_bit() noexcept + { + return 63; + } + static constexpr int minimum_exponent() noexcept + { + return -1023; + } + static constexpr int infinite_power() noexcept + { + return 0x7FF; + } + static constexpr int smallest_power_of_ten() noexcept + { + return -342; + } + static constexpr int largest_power_of_ten() noexcept + { + return 308; + } + static constexpr int min_exponent_round_to_even() noexcept + { + return -4; + } + static constexpr int max_exponent_round_to_even() noexcept + { + return 23; + } + static constexpr int max_exponent_fast_path() noexcept + { + return 22; + } + static constexpr std::uint64_t max_mantissa_fast_path() noexcept + { + return std::uint64_t{2} << 52u; + } + static constexpr std::int64_t max_digits() noexcept + { + return 769; + } +}; /*! -@brief parse a float with std::from_chars (Eisel-Lemire) when available +@brief whether @a FloatType is IEEE-754 binary32 or binary64 -std::from_chars is locale-independent, correctly rounded, and - via the -Eisel-Lemire algorithm in modern standard libraries - much faster than strtod -over the whole value range (not just the Clinger subset). It is used only when -__cpp_lib_to_chars indicates full floating-point support and only when it -consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so -the caller's strtod fallback supplies the well-defined ±inf/0 result the parser -expects (side-stepping the P4168 divergence between implementations). - -@return true if the value was parsed exactly and fully; false to fall back +These formats (float, double, and long double where it is binary64, e.g. +with MSVC or on Apple arm64) are converted by parse_float_native(). The +predicate is the one the serializer uses to choose Grisu2. */ template -bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept +struct has_native_float_format { - // JSON_HAS_CPP_17 must gate the use as well as the include above: - // some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even - // in C++14 mode, where is not included. -#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars) - const auto result = std::from_chars(first, last, out); - return result.ec == std::errc() && result.ptr == last; -#else - static_cast(first); - static_cast(last); - static_cast(out); - return false; -#endif -} + static constexpr bool value = + (std::numeric_limits::is_iec559 && std::numeric_limits::digits == 24 && std::numeric_limits::max_exponent == 128) || + (std::numeric_limits::is_iec559 && std::numeric_limits::digits == 53 && std::numeric_limits::max_exponent == 1024); +}; -/// whether the eight bytes of @a v (see read_eight_bytes()) are ASCII digits -/// (after fast_float's is_made_of_eight_digits_fast) -inline bool is_eight_digits(std::uint64_t v) noexcept -{ - return ((v & 0xF0F0F0F0F0F0F0F0u) | (((v + 0x0606060606060606u) & 0xF0F0F0F0F0F0F0F0u) >> 4u)) == 0x3333333333333333u; -} +/// the C++ type (float or double) that holds a binary32 or binary64 @a FloatType +template +using native_float_t = typename std::conditional::digits == 24, float, double>::type; /// the value of the eight ASCII digits in @a v (see read_eight_bytes()), three /// multiplications instead of eight (after simdjson and fast_float) @@ -317,31 +256,157 @@ inline std::uint32_t parse_eight_digits(std::uint64_t v) noexcept return static_cast(((v & 0x0000FFFF0000FFFFu) * 42949672960001u) >> 32u); } +/// whether [first, last) contains a digit other than '0' +inline bool has_nonzero_digit(const char* first, const char* last) noexcept +{ + for (; first != last; ++first) + { + if (*first != '0') + { + return true; + } + } + return false; +} + +/// the value of the validated exponent digits [+-]?[0-9]+ in [first, last), +/// saturated far beyond every range +inline std::int64_t parse_float_exponent(const char* first, const char* last) noexcept +{ + const bool negative = *first == '-'; + first += (*first == '-' || *first == '+') ? 1 : 0; + constexpr std::int64_t saturation = 100000000000000000; // 10^17 + std::int64_t value = 0; + for (; first != last; ++first) + { + if (value < saturation) + { + value = (value * 10) + (*first - '0'); + } + } + return negative ? -value : value; +} + +/// a float token as w * 10^exponent, see parse_float_significand() +struct float_significand +{ + std::uint64_t w = 0; ///< the first (at most 19) significant digits + std::int64_t exponent = 0; ///< the decimal exponent of the last digit in w + bool negative = false; ///< whether the token starts with '-' + bool truncated = false; ///< whether nonzero digits follow the ones in w +}; + /*! -@brief the double nearest to w * 10^q (Eisel-Lemire) +@brief split a validated number token into sign, significand, and exponent + +The lexer has validated the token against the JSON grammar and knows where its +parts are, so this needs no character classification: the integer part ends at +@a decimal_point_position (or @a mantissa_end), the fraction at @a mantissa_end, +and an exponent follows. At most 19 significant digits are kept; the value then +lies in [w, w + 1) * 10^exponent, and is exactly w * 10^exponent unless +truncated is set. + +@param[in] first pointer to the first character of the token +@param[in] last pointer past the last character +@param[in] decimal_point_position index of the '.' in the token, or + std::string::npos if there is none +@param[in] mantissa_end index of the 'e'/'E', or the token length +*/ +inline float_significand parse_float_significand(const char* first, const char* last, + std::size_t decimal_point_position, std::size_t mantissa_end) noexcept +{ + float_significand s; + const char* p = first; + s.negative = *p == '-'; + p += s.negative ? 1 : 0; + const bool has_dot = decimal_point_position != std::string::npos; + const char* const mantissa_last = first + mantissa_end; + const char* const integer_last = has_dot ? first + decimal_point_position : mantissa_last; + + std::uint64_t w = 0; + int remaining = 19; // digits that still fit into w + if (*p != '0') // the integer part is "0" or [1-9][0-9]* + { + while (remaining >= 8 && integer_last - p >= 8) + { + w = (w * 100000000u) + parse_eight_digits(read_eight_bytes(p)); + p += 8; + remaining -= 8; + } + for (; remaining > 0 && p != integer_last; ++p, --remaining) + { + w = (w * 10u) + static_cast(*p - '0'); + } + s.exponent = integer_last - p; + s.truncated = has_nonzero_digit(p, integer_last); + } + + if (has_dot) + { + p = integer_last + 1; + if (w == 0) + { + // zeros after the decimal point of "0." are not significant + const char* const zeros = p; + while (p != mantissa_last && *p == '0') + { + ++p; + } + s.exponent -= p - zeros; + } + const char* const digits = p; + while (remaining >= 8 && mantissa_last - p >= 8) + { + w = (w * 100000000u) + parse_eight_digits(read_eight_bytes(p)); + p += 8; + remaining -= 8; + } + for (; remaining > 0 && p != mantissa_last; ++p, --remaining) + { + w = (w * 10u) + static_cast(*p - '0'); + } + s.exponent -= p - digits; + s.truncated = s.truncated || has_nonzero_digit(p, mantissa_last); + } + + if (mantissa_last != last) + { + s.exponent += parse_float_exponent(mantissa_last + 1, last); + } + s.w = w; + return s; +} + +/*! +@brief the bits of the float nearest to w * 10^q (Eisel-Lemire) The algorithm of Daniel Lemire, "Number Parsing at a Gigabyte per Second" (Software: Practice and Experience, 2021), after fast_float's compute_float (used under the MIT license). With a 128-bit approximation of 5^q, the product -is always sufficient to round correctly for w with at most 19 digits (Noble -Mushtak and Daniel Lemire, "Fast number parsing without fallback", Software: -Practice and Experience, 2023). Only integer arithmetic is used, so the result -does not depend on the floating-point environment. +is always sufficient to round correctly for w < 2^64 (Noble Mushtak and Daniel +Lemire, "Fast number parsing without fallback", Software: Practice and +Experience, 2023). Only integer arithmetic is used, so the result does not +depend on the floating-point environment. +It is always inlined, like decimal_to_float(), so that hot loops of callers +keep the whole conversion inline. + +@tparam Format ieee_binary_format<24> (binary32) or ieee_binary_format<53> (binary64) @param[in] q decimal exponent -@param[in] w significand, w != 0 +@param[in] w significand @return the IEEE-754 bits of the positive result (0 for underflow, infinity for overflow) */ -inline std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept +template +JSON_HEDLEY_ALWAYS_INLINE std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept { - constexpr int mantissa_bits = 52; - constexpr std::uint64_t infinity = std::uint64_t{0x7FF} << mantissa_bits; - if (q < pow5_128_smallest_power) + constexpr int mantissa_bits = Format::mantissa_bits(); + constexpr std::uint64_t infinity = static_cast(Format::infinite_power()) << mantissa_bits; + if (w == 0 || q < Format::smallest_power_of_ten()) { return 0; } - if (q > pow5_128_largest_power) + if (q > Format::largest_power_of_ten()) { return infinity; } @@ -365,8 +430,8 @@ inline std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept const auto upperbit = static_cast(product.high >> 63u); const int shift = upperbit + 64 - mantissa_bits - 3; std::uint64_t mantissa = product.high >> static_cast(shift); - // floor(log2(10^q)) + 63 + 1023, with log2(10) ~ 217706 / 2^16 - std::int64_t power2 = (((152170 + 65536) * q) >> 16) + 63 + upperbit - lz + 1023; + // floor(log2(10^q)) + 63 + bias, with log2(10) ~ 217706 / 2^16 + std::int64_t power2 = (((152170 + 65536) * q) >> 16) + 63 + upperbit - lz - Format::minimum_exponent(); if (power2 <= 0) // subnormal { @@ -375,17 +440,18 @@ inline std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept return 0; } mantissa >>= static_cast(-power2 + 1); + // no tie is possible here: that needs a small |q| mantissa += (mantissa & 1u); mantissa >>= 1u; // rounding up may produce the smallest normal number power2 = (mantissa < (std::uint64_t{1} << mantissa_bits)) ? 0 : 1; - return mantissa | (static_cast(power2) << mantissa_bits); + return (mantissa & ((std::uint64_t{1} << mantissa_bits) - 1)) | (static_cast(power2) << mantissa_bits); } - // a value exactly between two doubles rounds to even; this can only + // a value exactly between two floats rounds to even; this can only // happen for small |q|, where 5^q is exact - if (product.low <= 1 && q >= -4 && q <= 23 && (mantissa & 3u) == 1 - && (mantissa << static_cast(shift)) == product.high) + if (product.low <= 1 && q >= Format::min_exponent_round_to_even() && q <= Format::max_exponent_round_to_even() + && (mantissa & 3u) == 1 && (mantissa << static_cast(shift)) == product.high) { mantissa &= ~std::uint64_t{1}; } @@ -397,198 +463,420 @@ inline std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept ++power2; } mantissa &= ~(std::uint64_t{1} << mantissa_bits); - if (power2 >= 0x7FF) + if (power2 >= Format::infinite_power()) { return infinity; } return mantissa | (static_cast(power2) << mantissa_bits); } -/*! -@brief parse a validated float token with the Eisel-Lemire algorithm - -The significand is accumulated eight digits at a time where possible. A token -with more than 19 significant digits is truncated to w; the value then lies -in [w, w + 1) * 10^q, and it is only returned if both ends round to the same -double, which covers all but a few such tokens. - -@param[in] first pointer to the first character of the token -@param[in] last pointer past the last character -@param[out] out the correctly rounded value on success (±infinity if it - overflows, like strtod) -@return true on success; false if strtod must decide -*/ -inline bool parse_float_eisel_lemire(const char* first, const char* last, double& out) noexcept +/// an unsigned integer of up to 4096 bits for digit_comparison() (32-bit limbs, +/// so only 32x32->64-bit multiplications are needed) +class float_bigint { - const char* p = first; - const bool negative = (p != last && *p == '-'); - if (negative) + public: + explicit float_bigint(std::uint64_t value) noexcept { - ++p; + for (; value != 0; value >>= 32u) + { + limbs[count++] = static_cast(value); + } } - std::uint64_t w = 0; - int digits = 0; // significant digits in w - std::int64_t exponent = 0; - bool truncated = false; - bool in_fraction = false; - for (;;) + /// *this = *this * factor + summand + void multiply_add(std::uint32_t factor, std::uint32_t summand) noexcept { - // eight digits at a time, as long as they fit into w - while (w != 0 && digits <= 19 - 8 && last - p >= 8) + std::uint64_t carry = summand; + for (std::size_t i = 0; i < count; ++i) { - const std::uint64_t v = read_eight_bytes(p); - if (!is_eight_digits(v)) - { - break; - } - w = (w * 100000000u) + parse_eight_digits(v); - digits += 8; - exponent -= in_fraction ? 8 : 0; - p += 8; + const std::uint64_t product = (static_cast(limbs[i]) * factor) + carry; + limbs[i] = static_cast(product); + carry = product >> 32u; } - if (p == last) + if (carry != 0) { - break; + JSON_ASSERT(count < limbs.size()); + limbs[count++] = static_cast(carry); } - const char c = *p; - if (c >= '0' && c <= '9') + } + + /// *this = *this * 5^n + void multiply_power_of_five(std::int64_t n) noexcept + { + static const std::array powers = { - if (w == 0 && c == '0') - { - // leading zeros are not significant, but scale a fraction - exponent -= in_fraction ? 1 : 0; - } - else if (digits < 19) - { - w = (w * 10u) + static_cast(c - '0'); - ++digits; - exponent -= in_fraction ? 1 : 0; - } - else - { - // dropped: the value lies between w and w + 1 (in units of - // the last kept digit) unless all dropped digits are zero - truncated = truncated || c != '0'; - exponent += in_fraction ? 0 : 1; - } - ++p; + {1u, 5u, 25u, 125u, 625u, 3125u, 15625u, 78125u, 390625u, 1953125u, 9765625u, 48828125u, 244140625u, 1220703125u} + }; + for (; n >= 13; n -= 13) + { + multiply_add(powers[13], 0); } - else if (c == '.') + multiply_add(powers[static_cast(n)], 0); + } + + /// *this = *this * 2^n + void shift_left(std::int64_t n) noexcept + { + if (count == 0) { - in_fraction = true; - ++p; + return; + } + const auto limb_shift = static_cast(n / 32); + const auto bit_shift = static_cast(n % 32); + JSON_ASSERT(count + limb_shift + 1 <= limbs.size()); + if (bit_shift != 0) + { + std::uint32_t carry = 0; + for (std::size_t i = 0; i < count; ++i) + { + const std::uint32_t limb = limbs[i]; + limbs[i] = (limb << bit_shift) | carry; + carry = limb >> (32u - bit_shift); + } + if (carry != 0) + { + limbs[count++] = carry; + } + } + if (limb_shift != 0) + { + for (std::size_t i = count; i-- > 0;) + { + limbs[i + limb_shift] = limbs[i]; + } + for (std::size_t i = 0; i < limb_shift; ++i) + { + limbs[i] = 0; + } + count += limb_shift; + } + } + + /// -1, 0, or 1 if *this is less than, equal to, or greater than @a other + int compare(const float_bigint& other) const noexcept + { + if (count != other.count) + { + return count < other.count ? -1 : 1; + } + for (std::size_t i = count; i-- > 0;) + { + if (limbs[i] != other.limbs[i]) + { + return limbs[i] < other.limbs[i] ? -1 : 1; + } + } + return 0; + } + + private: + std::array limbs{{}}; + std::size_t count = 0; +}; + +/*! +@brief round a token exactly when eisel_lemire() cannot decide (slow path) + +The value v of the token lies strictly between two adjacent floats, whose +lower one has the bits @a lower, and the result depends on whether v is below, +at, or above the midpoint m between them. Both are compared exactly as big +integers: v = D * 10^s with the significant digits D (at most +Format::max_digits() of them, more than any midpoint has; further nonzero +digits only put v above m) and m = (2 * mantissa + 1) * 2^(e - 1). This is the +digit comparison of fast_float (Daniel Lemire and contributors, used under the +MIT license), simplified by starting from the two candidates. + +@param[in] first pointer to the first character of the token +@param[in] last pointer past the last character +@param[in] lower the bits of the float below v +@return the bits of the correctly rounded result +*/ +template +std::uint64_t digit_comparison(const char* first, const char* last, std::uint64_t lower) noexcept +{ + const char* p = first + ((*first == '-') ? 1 : 0); + + // D, in chunks of up to 9 digits, and s + float_bigint digits(0); + std::int64_t count = 0; + std::int64_t point = 0; // the value is 0.D... * 10^point + bool truncated = false; + std::uint32_t chunk = 0; + int chunk_digits = 0; + static const std::array powers_of_ten = {{1u, 10u, 100u, 1000u, 10000u, 100000u, 1000000u, 10000000u, 100000000u, 1000000000u}}; + const auto append = [&](char c) noexcept + { + if (count < Format::max_digits()) + { + chunk = (chunk * 10u) + static_cast(c - '0'); + ++count; + if (++chunk_digits == 9) + { + digits.multiply_add(powers_of_ten[9], chunk); + chunk = 0; + chunk_digits = 0; + } } else { - break; // 'e' or 'E' + truncated = truncated || c != '0'; + } + }; + bool significant = false; + for (; p != last && *p >= '0' && *p <= '9'; ++p) + { + significant = significant || *p != '0'; + if (significant) + { + append(*p); + ++point; } } - - if (p != last) + if (p != last && *p == '.') { - ++p; // 'e' or 'E' - bool exp_negative = false; - if (p != last && (*p == '-' || *p == '+')) + for (++p; p != last && *p >= '0' && *p <= '9'; ++p) { - exp_negative = (*p == '-'); - ++p; - } - std::int64_t exp_value = 0; - for (; p != last; ++p) - { - // saturate: any exponent beyond this under- or overflows anyway - if (exp_value < 100000) + significant = significant || *p != '0'; + if (significant) { - exp_value = (exp_value * 10) + (*p - '0'); + append(*p); + } + else + { + --point; } } - exponent += exp_negative ? -exp_value : exp_value; } - - std::uint64_t bits = 0; - if (w != 0) + if (chunk_digits != 0) { - bits = eisel_lemire(exponent, w); - if (truncated && (w + 1 == 0 || eisel_lemire(exponent, w + 1) != bits)) - { - return false; - } + digits.multiply_add(powers_of_ten[static_cast(chunk_digits)], chunk); } - bits |= negative ? (std::uint64_t{1} << 63u) : 0u; - static_assert(sizeof(double) == sizeof(std::uint64_t), "double must have 64 bits"); - std::memcpy(&out, &bits, sizeof(out)); - return true; + if (p != last) + { + point += parse_float_exponent(p + 1, last); + } + const std::int64_t s = point - count; // v = D * 10^s + + // the midpoint above the lower candidate + constexpr int mantissa_bits = Format::mantissa_bits(); + const std::uint64_t exponent_field = lower >> mantissa_bits; + std::uint64_t mantissa = lower & ((std::uint64_t{1} << mantissa_bits) - 1); + std::int64_t e = 1 + Format::minimum_exponent() - mantissa_bits; // of the smallest subnormal number + if (exponent_field != 0) + { + mantissa |= std::uint64_t{1} << mantissa_bits; + e += static_cast(exponent_field) - 1; + } + float_bigint midpoint((2 * mantissa) + 1); + const std::int64_t midpoint_exponent = e - 1; // m = midpoint * 2^midpoint_exponent + + // compare D * 5^s * 2^s with midpoint * 2^midpoint_exponent + if (s >= 0) + { + digits.multiply_power_of_five(s); + } + else + { + midpoint.multiply_power_of_five(-s); + } + const std::int64_t shift = s - midpoint_exponent; + if (shift >= 0) + { + digits.shift_left(shift); + } + else + { + midpoint.shift_left(-shift); + } + const int order = digits.compare(midpoint); + const bool round_up = order > 0 || (order == 0 && (truncated || (mantissa & 1u) != 0)); + return lower + (round_up ? 1u : 0u); } -/// Eisel-Lemire is only implemented for `double` -template -bool parse_float_eisel_lemire(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept +/// the double with the IEEE-754 bits @a bits +inline void float_from_bits(std::uint64_t bits, double& value) noexcept { - return false; + static_assert(sizeof(double) == sizeof(std::uint64_t), "double must have 64 bits"); + std::memcpy(&value, &bits, sizeof(value)); +} + +/// the float with the IEEE-754 bits @a bits (the lower 32) +inline void float_from_bits(std::uint64_t bits, float& value) noexcept +{ + static_assert(sizeof(float) == sizeof(std::uint32_t), "float must have 32 bits"); + const auto bits32 = static_cast(bits); + std::memcpy(&value, &bits32, sizeof(value)); +} + +/// the powers of ten that are exact in binary64 (up to 10^22) +inline double exact_power_of_ten(std::int64_t n, double /*tag*/) noexcept +{ + static const std::array powers = + { + { + 1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11, + 1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22 + } + }; + return powers[static_cast(n)]; +} + +/// the powers of ten that are exact in binary32 (up to 10^10) +inline float exact_power_of_ten(std::int64_t n, float /*tag*/) noexcept +{ + static const std::array powers = + { + {1e0f, 1e1f, 1e2f, 1e3f, 1e4f, 1e5f, 1e6f, 1e7f, 1e8f, 1e9f, 1e10f} + }; + return powers[static_cast(n)]; } /*! -@brief check whether Clinger's fast path can still succeed for a float token +@brief the binary32/binary64 value of a significand that was not truncated -parse_float_fast() needs a significand below 2^53. A mantissa with 17 or -more significant digits is at least 10^16 and therefore always exceeds it, -so calling the fast path would walk the token one extra time only to -decline before strtod has to run anyway. +The result is (-1)^negative * w * 10^exponent, correctly rounded (ties to +even): with Clinger's fast path where w and 10^|exponent| are exact, so that a +single floating-point operation rounds (only where intermediate results are not +kept in extended precision, see FLT_EVAL_METHOD), and with eisel_lemire() +otherwise. A value too large for the type becomes ±infinity, a value too small +±0. -Significant digits are the mantissa's digits from the first nonzero one on; -the sign, the decimal point, leading zeros, and the exponent do not count. -The answer is derived from indices - the digits are not scanned again - so -this stays off the hot path of the number scanners. +This is the core of the conversion that other parsers of JSON text share: they +can split a token themselves and still get the lexer's result. It is always +inlined, so that their hot loops keep the whole conversion inline. -@param[in] token the validated number token ('.' as decimal point) -@param[in] decimal_point_position index of the '.' in @a token, or - std::string::npos if there is none -@param[in] mantissa_end offset just past the last mantissa byte -@return false if parse_float_fast() is guaranteed to decline +@param[in] s the significand, with s.truncated == false */ -inline bool mantissa_fits_clinger(const char* token, std::size_t decimal_point_position, std::size_t mantissa_end) noexcept +template +JSON_HEDLEY_ALWAYS_INLINE FloatType decimal_to_float(const float_significand& s) noexcept { - // 10^16 already exceeds 2^53, so 17 digits can never fit - constexpr std::size_t limit = 17; + using result_type = native_float_t; + using format = ieee_binary_format::digits>; + JSON_ASSERT(!s.truncated); - const std::size_t neg = (token[0] == '-') ? 1u : 0u; - const std::size_t has_dot = (decimal_point_position != std::string::npos) ? 1u : 0u; - // the JSON grammar restricts the integer part to "0" or [1-9][0-9]*, so - // a leading zero can only be a lone "0", which is not significant - const std::size_t lead_zero = (token[neg] == '0') ? 1u : 0u; - JSON_ASSERT(mantissa_end >= neg + has_dot + lead_zero); - std::size_t digits = mantissa_end - neg - has_dot - lead_zero; - - if (JSON_HEDLEY_LIKELY(digits < limit)) +#if !defined(FLT_EVAL_METHOD) || FLT_EVAL_METHOD == 0 + if (s.exponent >= -format::max_exponent_fast_path() && s.exponent <= format::max_exponent_fast_path() + && s.w <= format::max_mantissa_fast_path()) { - return true; - } - - // Only a number below 1 can carry further insignificant zeros, and only - // while the count stays at the limit does removing them change the - // answer - so this loop is skipped for all but a few tokens. The - // fraction is located through decimal_point_position rather than by - // searching '.'. - if (lead_zero != 0) - { - JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit - for (std::size_t i = decimal_point_position + 1; - digits >= limit && i < mantissa_end && token[i] == '0'; ++i) + auto value = static_cast(s.w); + if (s.exponent < 0) { - --digits; + value /= exact_power_of_ten(-s.exponent, result_type{}); } + else + { + value *= exact_power_of_ten(s.exponent, result_type{}); + } + const FloatType result = s.negative ? -value : value; + return result; + } +#endif + + result_type value{}; + float_from_bits(eisel_lemire(s.exponent, s.w) | (s.negative ? (std::uint64_t{1} << format::sign_bit()) : 0u), value); + const FloatType result = value; + return result; +} + +/*! +@brief convert a validated number token to the nearest binary32/binary64 value + +The conversion is correctly rounded (ties to even) and independent of the +locale and of the C and C++ libraries: +1. parse_float_significand() splits the token into w * 10^q. +2. If no digits were dropped, decimal_to_float() rounds w * 10^q (Clinger's + fast path or Eisel-Lemire). +3. Otherwise, the value lies in [w, w + 1) * 10^q: if eisel_lemire() rounds + both ends to the same value, so does the token (in all but rare cases). +4. Otherwise, digit_comparison() compares the token exactly with the midpoint + between the two candidates. +A value too large for the type becomes ±infinity (the parser reports +out_of_range.406), a value too small ±0. + +@param[in] first pointer to the first character of the token +@param[in] last pointer past the last character +@param[in] decimal_point_position index of the '.' in the token, or + std::string::npos if there is none +@param[in] mantissa_end index of the 'e'/'E', or the token length +*/ +template +FloatType parse_float_native(const char* first, const char* last, + std::size_t decimal_point_position, std::size_t mantissa_end) noexcept +{ + using result_type = native_float_t; + using format = ieee_binary_format::digits>; + static_assert(std::numeric_limits::digits == std::numeric_limits::digits, "unexpected float format"); + + const float_significand s = parse_float_significand(first, last, decimal_point_position, mantissa_end); + if (JSON_HEDLEY_LIKELY(!s.truncated)) + { + return decimal_to_float(s); } - return digits < limit; + std::uint64_t bits = eisel_lemire(s.exponent, s.w); + if (JSON_HEDLEY_UNLIKELY(bits != eisel_lemire(s.exponent, s.w + 1))) + { + bits = digit_comparison(first, last, bits); + } + result_type value{}; + float_from_bits(bits | (s.negative ? (std::uint64_t{1} << format::sign_bit()) : 0u), value); + const FloatType result = value; + return result; +} + +/*! +@brief parse a float with std::from_chars when available + +Only used for the formats parse_float_native() does not convert (long double +formats other than binary64). std::from_chars is locale-independent and +correctly rounded. It is used only when __cpp_lib_to_chars indicates full +floating-point support and only when it consumes the entire token ([first, +last)). An under-/overflow (result_out_of_range) also declines, so the +caller's strtold fallback supplies the well-defined ±inf/0 result the parser +expects (side-stepping the P4168 divergence between implementations). + +@return true if the value was parsed exactly and fully; false to fall back +*/ +template +bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept +{ + // JSON_HAS_CPP_17 must gate the use as well as the include above: + // some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even + // in C++14 mode, where is not included. +#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars) + const auto result = std::from_chars(first, last, out); + return result.ec == std::errc() && result.ptr == last; +#else + static_cast(first); + static_cast(last); + static_cast(out); + return false; +#endif +} + +/// binary32 and binary64: the library's own conversion, which always succeeds +template +bool convert_float_fast(const char* first, const char* last, std::size_t decimal_point_position, + std::size_t mantissa_end, FloatType& value, std::true_type /*native*/) noexcept +{ + value = parse_float_native(first, last, decimal_point_position, mantissa_end); + return true; +} + +/// other formats (long double on x87, binary128, double-double): std::from_chars, if available +template +bool convert_float_fast(const char* first, const char* last, std::size_t /*decimal_point_position*/, + std::size_t /*mantissa_end*/, FloatType& value, std::false_type /*native*/) noexcept +{ + return parse_float_from_chars(first, last, value); } /*! @brief convert a validated float token without the C library, if possible -Tries std::from_chars (when available), Clinger's exact fast path (double -only, skipped when it cannot succeed), and the Eisel-Lemire algorithm (double -only). +float, double, and long double where it is binary64 are always converted, by +parse_float_native(). Other long double formats are converted with +std::from_chars where the standard library supports it. @param[in] first pointer to the first character of the token @param[in] last pointer past the last character @@ -604,19 +892,8 @@ template bool convert_float_fast(const char* first, const char* last, std::size_t decimal_point_position, std::size_t mantissa_end, FloatType& value) noexcept { - if (parse_float_from_chars(first, last, value)) - { - return true; - } - // Skipping a fast path that cannot succeed is lossless and saves a full - // extra pass over the token's bytes, which otherwise shows up on - // high-precision inputs such as canada.json - if (mantissa_fits_clinger(first, decimal_point_position, mantissa_end) - && parse_float_fast(first, last, value)) - { - return true; - } - return parse_float_eisel_lemire(first, last, value); + return convert_float_fast(first, last, decimal_point_position, mantissa_end, value, + std::integral_constant::value> {}); } /// std::strtof, std::strtod, or std::strtold, chosen by the type of @a f @@ -651,6 +928,11 @@ inline char get_decimal_point() noexcept /*! @brief convert a validated float token with strtof/strtod/strtold +Only used for what convert_float_fast() does not convert: long double formats +other than binary64 where std::from_chars is unavailable or reports an under- +or overflow, and floating-point types that are not IEEE-754 (see +has_native_float_format). + These functions expect the decimal point of the *current* locale, so it is looked up right before the conversion instead of once when the lexer is constructed: a locale change in between (by a parser callback, a SAX @@ -711,5 +993,33 @@ void convert_float_locale_aware(StringType& token, std::size_t decimal_point_pos } } +/*! +@brief convert a validated float token like the lexer does + +For parsers of JSON text other than the lexer, which converts its own token +buffer in place. float, double, and long double where it is binary64 are +converted without allocation and independent of the locale; only other long +double formats that std::from_chars does not support need a copy of the token +for convert_float_locale_aware(). + +@param[in] first pointer to the first character of the token +@param[in] last pointer past the last character +@param[in] decimal_point_position index of the '.' in the token, or + std::string::npos if there is none +@param[in] mantissa_end index of the 'e'/'E', or the token length +@return the value, ±infinity if it overflows +*/ +template +FloatType convert_float(const char* first, const char* last, std::size_t decimal_point_position, std::size_t mantissa_end) +{ + FloatType value{}; + if (!convert_float_fast(first, last, decimal_point_position, mantissa_end, value)) + { + std::string token(first, last); + convert_float_locale_aware(token, decimal_point_position, value); + } + return value; +} + } // namespace detail NLOHMANN_JSON_NAMESPACE_END diff --git a/single_include/nlohmann/json.hpp b/single_include/nlohmann/json.hpp index 1232e184c..7b43d5e5b 100644 --- a/single_include/nlohmann/json.hpp +++ b/single_include/nlohmann/json.hpp @@ -8502,6 +8502,7 @@ NLOHMANN_JSON_NAMESPACE_END #include // memcpy #include // numeric_limits #include // string +#include // conditional, integral_constant, true_type, false_type // #include // __ _____ _____ _____ @@ -8976,10 +8977,13 @@ NLOHMANN_JSON_NAMESPACE_END #endif // This file contains the value-conversion helpers used by the lexer to turn an -// already-validated number token into a value, without the locale/errno -// overhead of std::strtoull/std::strtod where possible. They are free functions -// so the lexer stays focused on scanning (see lexer::convert_number()) and so -// that other parsers of JSON text can convert tokens exactly like it does. +// already-validated number token into a value. Integers and binary32/binary64 +// floats (float, double, and long double where it is binary64) are converted +// by the library itself, without the locale/errno overhead of +// std::strtoull/std::strtod and correctly rounded; other long double formats +// use std::from_chars or std::strtold. They are free functions so the lexer +// stays focused on scanning (see lexer::convert_number()) and so that other +// parsers of JSON text can convert tokens exactly like it does. NLOHMANN_JSON_NAMESPACE_BEGIN namespace detail @@ -9056,198 +9060,133 @@ bool parse_integer_signed(const char* first, const char* last, NumberIntegerType } /*! -@brief exact fast path for parsing a `double` (Clinger's algorithm) - -For the common case - at most 19 significant digits, a decimal exponent in -[-22, 22], and a significand below 2^53 - the value equals significand * -10^exp computed in IEEE-754 double arithmetic, which is exact under -round-to-nearest because both operands are exactly representable. This is the -same fast path used by fast_float/simdjson; the general cases are left to -std::strtod. The parser only activates for number_float_t == double; float and -long double keep the std::strtof/std::strtold paths (see the templated overload -below). - -@param[in] first pointer to the first character of the number -@param[in] last pointer past the last character -@param[out] out the parsed value on success -@return true if the value was parsed exactly; false to fall back to strtod +@brief parameters of the IEEE-754 binary32 and binary64 formats for the float + conversion (after fast_float's binary_format) */ -inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept +template +struct ieee_binary_format; + +template<> +struct ieee_binary_format<24> // binary32 { -#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0 - // Clinger's fast path is only exact when double operations are evaluated in - // true double precision. On platforms that keep intermediates in extended - // precision (e.g. the x87 FPU on 32-bit x86, where FLT_EVAL_METHOD == 2) the - // single significand * 10^scale step is double-rounded and can be 1 ULP off, - // so decline and let the caller fall back to the correctly-rounded - // std::from_chars / std::strtod path. - static_cast(first); - static_cast(last); - static_cast(out); - return false; -#else - static const std::array powers_of_ten = + static constexpr int mantissa_bits() noexcept { - { - 1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11, - 1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22 - } - }; + return 23; + } + static constexpr int sign_bit() noexcept + { + return 31; + } + static constexpr int minimum_exponent() noexcept + { + return -127; + } + static constexpr int infinite_power() noexcept + { + return 0xFF; + } + // w * 10^q with w < 2^64 is below half the smallest subnormal number for + // q < smallest_power_of_ten() and at least infinity for q > largest_power_of_ten() + static constexpr int smallest_power_of_ten() noexcept + { + return -64; + } + static constexpr int largest_power_of_ten() noexcept + { + return 38; + } + // w * 10^q can only be exactly between two numbers for q in this range + static constexpr int min_exponent_round_to_even() noexcept + { + return -17; + } + static constexpr int max_exponent_round_to_even() noexcept + { + return 10; + } + // Clinger's fast path: w and 10^|q| are exact + static constexpr int max_exponent_fast_path() noexcept + { + return 10; + } + static constexpr std::uint64_t max_mantissa_fast_path() noexcept + { + return std::uint64_t{2} << 23u; + } + // a midpoint between two numbers has at most this many significant digits + static constexpr std::int64_t max_digits() noexcept + { + return 114; + } +}; - const char* p = first; - bool negative = false; - if (p != last && (*p == '-' || *p == '+')) - { - negative = (*p == '-'); - ++p; - } - - std::uint64_t significand = 0; - int num_digits = 0; - int fractional_digits = 0; - bool seen_dot = false; - bool any_digit = false; - for (; p != last; ++p) - { - const char c = *p; - if (c >= '0' && c <= '9') - { - any_digit = true; - if (JSON_HEDLEY_UNLIKELY(num_digits >= 19)) - { - return false; // significand may not fit into uint64_t - } - significand = (significand * 10u) + static_cast(c - '0'); - ++num_digits; - fractional_digits += static_cast(seen_dot); - } - else if (c == '.') - { - if (JSON_HEDLEY_UNLIKELY(seen_dot)) - { - return false; - } - seen_dot = true; - } - else if (c == 'e' || c == 'E') - { - ++p; - break; - } - else - { - return false; - } - } - if (JSON_HEDLEY_UNLIKELY(!any_digit)) - { - return false; - } - - int exponent = 0; - if (p != last) // an exponent part remains - { - bool exp_negative = false; - if (p != last && (*p == '-' || *p == '+')) - { - exp_negative = (*p == '-'); - ++p; - } - bool any_exp_digit = false; - for (; p != last; ++p) - { - if (JSON_HEDLEY_UNLIKELY(*p < '0' || *p > '9')) - { - return false; - } - exponent = (exponent * 10) + (*p - '0'); - any_exp_digit = true; - if (JSON_HEDLEY_UNLIKELY(exponent > 9999)) - { - return false; - } - } - if (JSON_HEDLEY_UNLIKELY(!any_exp_digit)) - { - return false; - } - if (exp_negative) - { - exponent = -exponent; - } - } - - const int scale = exponent - fractional_digits; - if (JSON_HEDLEY_UNLIKELY(significand >= (static_cast(1) << 53))) - { - return false; // significand not exactly representable as double - } - - auto result = static_cast(significand); - if (scale >= 0) - { - if (JSON_HEDLEY_UNLIKELY(scale > 22)) - { - return false; - } - result *= powers_of_ten[static_cast(scale)]; - } - else - { - if (JSON_HEDLEY_UNLIKELY(-scale > 22)) - { - return false; - } - result /= powers_of_ten[static_cast(-scale)]; - } - out = negative ? -result : result; - return true; -#endif -} - -/// fast float path is only exact for `double`; decline for float/long double -template -bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept +template<> +struct ieee_binary_format<53> // binary64 { - return false; -} + static constexpr int mantissa_bits() noexcept + { + return 52; + } + static constexpr int sign_bit() noexcept + { + return 63; + } + static constexpr int minimum_exponent() noexcept + { + return -1023; + } + static constexpr int infinite_power() noexcept + { + return 0x7FF; + } + static constexpr int smallest_power_of_ten() noexcept + { + return -342; + } + static constexpr int largest_power_of_ten() noexcept + { + return 308; + } + static constexpr int min_exponent_round_to_even() noexcept + { + return -4; + } + static constexpr int max_exponent_round_to_even() noexcept + { + return 23; + } + static constexpr int max_exponent_fast_path() noexcept + { + return 22; + } + static constexpr std::uint64_t max_mantissa_fast_path() noexcept + { + return std::uint64_t{2} << 52u; + } + static constexpr std::int64_t max_digits() noexcept + { + return 769; + } +}; /*! -@brief parse a float with std::from_chars (Eisel-Lemire) when available +@brief whether @a FloatType is IEEE-754 binary32 or binary64 -std::from_chars is locale-independent, correctly rounded, and - via the -Eisel-Lemire algorithm in modern standard libraries - much faster than strtod -over the whole value range (not just the Clinger subset). It is used only when -__cpp_lib_to_chars indicates full floating-point support and only when it -consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so -the caller's strtod fallback supplies the well-defined ±inf/0 result the parser -expects (side-stepping the P4168 divergence between implementations). - -@return true if the value was parsed exactly and fully; false to fall back +These formats (float, double, and long double where it is binary64, e.g. +with MSVC or on Apple arm64) are converted by parse_float_native(). The +predicate is the one the serializer uses to choose Grisu2. */ template -bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept +struct has_native_float_format { - // JSON_HAS_CPP_17 must gate the use as well as the include above: - // some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even - // in C++14 mode, where is not included. -#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars) - const auto result = std::from_chars(first, last, out); - return result.ec == std::errc() && result.ptr == last; -#else - static_cast(first); - static_cast(last); - static_cast(out); - return false; -#endif -} + static constexpr bool value = + (std::numeric_limits::is_iec559 && std::numeric_limits::digits == 24 && std::numeric_limits::max_exponent == 128) || + (std::numeric_limits::is_iec559 && std::numeric_limits::digits == 53 && std::numeric_limits::max_exponent == 1024); +}; -/// whether the eight bytes of @a v (see read_eight_bytes()) are ASCII digits -/// (after fast_float's is_made_of_eight_digits_fast) -inline bool is_eight_digits(std::uint64_t v) noexcept -{ - return ((v & 0xF0F0F0F0F0F0F0F0u) | (((v + 0x0606060606060606u) & 0xF0F0F0F0F0F0F0F0u) >> 4u)) == 0x3333333333333333u; -} +/// the C++ type (float or double) that holds a binary32 or binary64 @a FloatType +template +using native_float_t = typename std::conditional::digits == 24, float, double>::type; /// the value of the eight ASCII digits in @a v (see read_eight_bytes()), three /// multiplications instead of eight (after simdjson and fast_float) @@ -9258,31 +9197,157 @@ inline std::uint32_t parse_eight_digits(std::uint64_t v) noexcept return static_cast(((v & 0x0000FFFF0000FFFFu) * 42949672960001u) >> 32u); } +/// whether [first, last) contains a digit other than '0' +inline bool has_nonzero_digit(const char* first, const char* last) noexcept +{ + for (; first != last; ++first) + { + if (*first != '0') + { + return true; + } + } + return false; +} + +/// the value of the validated exponent digits [+-]?[0-9]+ in [first, last), +/// saturated far beyond every range +inline std::int64_t parse_float_exponent(const char* first, const char* last) noexcept +{ + const bool negative = *first == '-'; + first += (*first == '-' || *first == '+') ? 1 : 0; + constexpr std::int64_t saturation = 100000000000000000; // 10^17 + std::int64_t value = 0; + for (; first != last; ++first) + { + if (value < saturation) + { + value = (value * 10) + (*first - '0'); + } + } + return negative ? -value : value; +} + +/// a float token as w * 10^exponent, see parse_float_significand() +struct float_significand +{ + std::uint64_t w = 0; ///< the first (at most 19) significant digits + std::int64_t exponent = 0; ///< the decimal exponent of the last digit in w + bool negative = false; ///< whether the token starts with '-' + bool truncated = false; ///< whether nonzero digits follow the ones in w +}; + /*! -@brief the double nearest to w * 10^q (Eisel-Lemire) +@brief split a validated number token into sign, significand, and exponent + +The lexer has validated the token against the JSON grammar and knows where its +parts are, so this needs no character classification: the integer part ends at +@a decimal_point_position (or @a mantissa_end), the fraction at @a mantissa_end, +and an exponent follows. At most 19 significant digits are kept; the value then +lies in [w, w + 1) * 10^exponent, and is exactly w * 10^exponent unless +truncated is set. + +@param[in] first pointer to the first character of the token +@param[in] last pointer past the last character +@param[in] decimal_point_position index of the '.' in the token, or + std::string::npos if there is none +@param[in] mantissa_end index of the 'e'/'E', or the token length +*/ +inline float_significand parse_float_significand(const char* first, const char* last, + std::size_t decimal_point_position, std::size_t mantissa_end) noexcept +{ + float_significand s; + const char* p = first; + s.negative = *p == '-'; + p += s.negative ? 1 : 0; + const bool has_dot = decimal_point_position != std::string::npos; + const char* const mantissa_last = first + mantissa_end; + const char* const integer_last = has_dot ? first + decimal_point_position : mantissa_last; + + std::uint64_t w = 0; + int remaining = 19; // digits that still fit into w + if (*p != '0') // the integer part is "0" or [1-9][0-9]* + { + while (remaining >= 8 && integer_last - p >= 8) + { + w = (w * 100000000u) + parse_eight_digits(read_eight_bytes(p)); + p += 8; + remaining -= 8; + } + for (; remaining > 0 && p != integer_last; ++p, --remaining) + { + w = (w * 10u) + static_cast(*p - '0'); + } + s.exponent = integer_last - p; + s.truncated = has_nonzero_digit(p, integer_last); + } + + if (has_dot) + { + p = integer_last + 1; + if (w == 0) + { + // zeros after the decimal point of "0." are not significant + const char* const zeros = p; + while (p != mantissa_last && *p == '0') + { + ++p; + } + s.exponent -= p - zeros; + } + const char* const digits = p; + while (remaining >= 8 && mantissa_last - p >= 8) + { + w = (w * 100000000u) + parse_eight_digits(read_eight_bytes(p)); + p += 8; + remaining -= 8; + } + for (; remaining > 0 && p != mantissa_last; ++p, --remaining) + { + w = (w * 10u) + static_cast(*p - '0'); + } + s.exponent -= p - digits; + s.truncated = s.truncated || has_nonzero_digit(p, mantissa_last); + } + + if (mantissa_last != last) + { + s.exponent += parse_float_exponent(mantissa_last + 1, last); + } + s.w = w; + return s; +} + +/*! +@brief the bits of the float nearest to w * 10^q (Eisel-Lemire) The algorithm of Daniel Lemire, "Number Parsing at a Gigabyte per Second" (Software: Practice and Experience, 2021), after fast_float's compute_float (used under the MIT license). With a 128-bit approximation of 5^q, the product -is always sufficient to round correctly for w with at most 19 digits (Noble -Mushtak and Daniel Lemire, "Fast number parsing without fallback", Software: -Practice and Experience, 2023). Only integer arithmetic is used, so the result -does not depend on the floating-point environment. +is always sufficient to round correctly for w < 2^64 (Noble Mushtak and Daniel +Lemire, "Fast number parsing without fallback", Software: Practice and +Experience, 2023). Only integer arithmetic is used, so the result does not +depend on the floating-point environment. +It is always inlined, like decimal_to_float(), so that hot loops of callers +keep the whole conversion inline. + +@tparam Format ieee_binary_format<24> (binary32) or ieee_binary_format<53> (binary64) @param[in] q decimal exponent -@param[in] w significand, w != 0 +@param[in] w significand @return the IEEE-754 bits of the positive result (0 for underflow, infinity for overflow) */ -inline std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept +template +JSON_HEDLEY_ALWAYS_INLINE std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept { - constexpr int mantissa_bits = 52; - constexpr std::uint64_t infinity = std::uint64_t{0x7FF} << mantissa_bits; - if (q < pow5_128_smallest_power) + constexpr int mantissa_bits = Format::mantissa_bits(); + constexpr std::uint64_t infinity = static_cast(Format::infinite_power()) << mantissa_bits; + if (w == 0 || q < Format::smallest_power_of_ten()) { return 0; } - if (q > pow5_128_largest_power) + if (q > Format::largest_power_of_ten()) { return infinity; } @@ -9306,8 +9371,8 @@ inline std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept const auto upperbit = static_cast(product.high >> 63u); const int shift = upperbit + 64 - mantissa_bits - 3; std::uint64_t mantissa = product.high >> static_cast(shift); - // floor(log2(10^q)) + 63 + 1023, with log2(10) ~ 217706 / 2^16 - std::int64_t power2 = (((152170 + 65536) * q) >> 16) + 63 + upperbit - lz + 1023; + // floor(log2(10^q)) + 63 + bias, with log2(10) ~ 217706 / 2^16 + std::int64_t power2 = (((152170 + 65536) * q) >> 16) + 63 + upperbit - lz - Format::minimum_exponent(); if (power2 <= 0) // subnormal { @@ -9316,17 +9381,18 @@ inline std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept return 0; } mantissa >>= static_cast(-power2 + 1); + // no tie is possible here: that needs a small |q| mantissa += (mantissa & 1u); mantissa >>= 1u; // rounding up may produce the smallest normal number power2 = (mantissa < (std::uint64_t{1} << mantissa_bits)) ? 0 : 1; - return mantissa | (static_cast(power2) << mantissa_bits); + return (mantissa & ((std::uint64_t{1} << mantissa_bits) - 1)) | (static_cast(power2) << mantissa_bits); } - // a value exactly between two doubles rounds to even; this can only + // a value exactly between two floats rounds to even; this can only // happen for small |q|, where 5^q is exact - if (product.low <= 1 && q >= -4 && q <= 23 && (mantissa & 3u) == 1 - && (mantissa << static_cast(shift)) == product.high) + if (product.low <= 1 && q >= Format::min_exponent_round_to_even() && q <= Format::max_exponent_round_to_even() + && (mantissa & 3u) == 1 && (mantissa << static_cast(shift)) == product.high) { mantissa &= ~std::uint64_t{1}; } @@ -9338,198 +9404,420 @@ inline std::uint64_t eisel_lemire(std::int64_t q, std::uint64_t w) noexcept ++power2; } mantissa &= ~(std::uint64_t{1} << mantissa_bits); - if (power2 >= 0x7FF) + if (power2 >= Format::infinite_power()) { return infinity; } return mantissa | (static_cast(power2) << mantissa_bits); } -/*! -@brief parse a validated float token with the Eisel-Lemire algorithm - -The significand is accumulated eight digits at a time where possible. A token -with more than 19 significant digits is truncated to w; the value then lies -in [w, w + 1) * 10^q, and it is only returned if both ends round to the same -double, which covers all but a few such tokens. - -@param[in] first pointer to the first character of the token -@param[in] last pointer past the last character -@param[out] out the correctly rounded value on success (±infinity if it - overflows, like strtod) -@return true on success; false if strtod must decide -*/ -inline bool parse_float_eisel_lemire(const char* first, const char* last, double& out) noexcept +/// an unsigned integer of up to 4096 bits for digit_comparison() (32-bit limbs, +/// so only 32x32->64-bit multiplications are needed) +class float_bigint { - const char* p = first; - const bool negative = (p != last && *p == '-'); - if (negative) + public: + explicit float_bigint(std::uint64_t value) noexcept { - ++p; + for (; value != 0; value >>= 32u) + { + limbs[count++] = static_cast(value); + } } - std::uint64_t w = 0; - int digits = 0; // significant digits in w - std::int64_t exponent = 0; - bool truncated = false; - bool in_fraction = false; - for (;;) + /// *this = *this * factor + summand + void multiply_add(std::uint32_t factor, std::uint32_t summand) noexcept { - // eight digits at a time, as long as they fit into w - while (w != 0 && digits <= 19 - 8 && last - p >= 8) + std::uint64_t carry = summand; + for (std::size_t i = 0; i < count; ++i) { - const std::uint64_t v = read_eight_bytes(p); - if (!is_eight_digits(v)) - { - break; - } - w = (w * 100000000u) + parse_eight_digits(v); - digits += 8; - exponent -= in_fraction ? 8 : 0; - p += 8; + const std::uint64_t product = (static_cast(limbs[i]) * factor) + carry; + limbs[i] = static_cast(product); + carry = product >> 32u; } - if (p == last) + if (carry != 0) { - break; + JSON_ASSERT(count < limbs.size()); + limbs[count++] = static_cast(carry); } - const char c = *p; - if (c >= '0' && c <= '9') + } + + /// *this = *this * 5^n + void multiply_power_of_five(std::int64_t n) noexcept + { + static const std::array powers = { - if (w == 0 && c == '0') - { - // leading zeros are not significant, but scale a fraction - exponent -= in_fraction ? 1 : 0; - } - else if (digits < 19) - { - w = (w * 10u) + static_cast(c - '0'); - ++digits; - exponent -= in_fraction ? 1 : 0; - } - else - { - // dropped: the value lies between w and w + 1 (in units of - // the last kept digit) unless all dropped digits are zero - truncated = truncated || c != '0'; - exponent += in_fraction ? 0 : 1; - } - ++p; + {1u, 5u, 25u, 125u, 625u, 3125u, 15625u, 78125u, 390625u, 1953125u, 9765625u, 48828125u, 244140625u, 1220703125u} + }; + for (; n >= 13; n -= 13) + { + multiply_add(powers[13], 0); } - else if (c == '.') + multiply_add(powers[static_cast(n)], 0); + } + + /// *this = *this * 2^n + void shift_left(std::int64_t n) noexcept + { + if (count == 0) { - in_fraction = true; - ++p; + return; + } + const auto limb_shift = static_cast(n / 32); + const auto bit_shift = static_cast(n % 32); + JSON_ASSERT(count + limb_shift + 1 <= limbs.size()); + if (bit_shift != 0) + { + std::uint32_t carry = 0; + for (std::size_t i = 0; i < count; ++i) + { + const std::uint32_t limb = limbs[i]; + limbs[i] = (limb << bit_shift) | carry; + carry = limb >> (32u - bit_shift); + } + if (carry != 0) + { + limbs[count++] = carry; + } + } + if (limb_shift != 0) + { + for (std::size_t i = count; i-- > 0;) + { + limbs[i + limb_shift] = limbs[i]; + } + for (std::size_t i = 0; i < limb_shift; ++i) + { + limbs[i] = 0; + } + count += limb_shift; + } + } + + /// -1, 0, or 1 if *this is less than, equal to, or greater than @a other + int compare(const float_bigint& other) const noexcept + { + if (count != other.count) + { + return count < other.count ? -1 : 1; + } + for (std::size_t i = count; i-- > 0;) + { + if (limbs[i] != other.limbs[i]) + { + return limbs[i] < other.limbs[i] ? -1 : 1; + } + } + return 0; + } + + private: + std::array limbs{{}}; + std::size_t count = 0; +}; + +/*! +@brief round a token exactly when eisel_lemire() cannot decide (slow path) + +The value v of the token lies strictly between two adjacent floats, whose +lower one has the bits @a lower, and the result depends on whether v is below, +at, or above the midpoint m between them. Both are compared exactly as big +integers: v = D * 10^s with the significant digits D (at most +Format::max_digits() of them, more than any midpoint has; further nonzero +digits only put v above m) and m = (2 * mantissa + 1) * 2^(e - 1). This is the +digit comparison of fast_float (Daniel Lemire and contributors, used under the +MIT license), simplified by starting from the two candidates. + +@param[in] first pointer to the first character of the token +@param[in] last pointer past the last character +@param[in] lower the bits of the float below v +@return the bits of the correctly rounded result +*/ +template +std::uint64_t digit_comparison(const char* first, const char* last, std::uint64_t lower) noexcept +{ + const char* p = first + ((*first == '-') ? 1 : 0); + + // D, in chunks of up to 9 digits, and s + float_bigint digits(0); + std::int64_t count = 0; + std::int64_t point = 0; // the value is 0.D... * 10^point + bool truncated = false; + std::uint32_t chunk = 0; + int chunk_digits = 0; + static const std::array powers_of_ten = {{1u, 10u, 100u, 1000u, 10000u, 100000u, 1000000u, 10000000u, 100000000u, 1000000000u}}; + const auto append = [&](char c) noexcept + { + if (count < Format::max_digits()) + { + chunk = (chunk * 10u) + static_cast(c - '0'); + ++count; + if (++chunk_digits == 9) + { + digits.multiply_add(powers_of_ten[9], chunk); + chunk = 0; + chunk_digits = 0; + } } else { - break; // 'e' or 'E' + truncated = truncated || c != '0'; + } + }; + bool significant = false; + for (; p != last && *p >= '0' && *p <= '9'; ++p) + { + significant = significant || *p != '0'; + if (significant) + { + append(*p); + ++point; } } - - if (p != last) + if (p != last && *p == '.') { - ++p; // 'e' or 'E' - bool exp_negative = false; - if (p != last && (*p == '-' || *p == '+')) + for (++p; p != last && *p >= '0' && *p <= '9'; ++p) { - exp_negative = (*p == '-'); - ++p; - } - std::int64_t exp_value = 0; - for (; p != last; ++p) - { - // saturate: any exponent beyond this under- or overflows anyway - if (exp_value < 100000) + significant = significant || *p != '0'; + if (significant) { - exp_value = (exp_value * 10) + (*p - '0'); + append(*p); + } + else + { + --point; } } - exponent += exp_negative ? -exp_value : exp_value; } - - std::uint64_t bits = 0; - if (w != 0) + if (chunk_digits != 0) { - bits = eisel_lemire(exponent, w); - if (truncated && (w + 1 == 0 || eisel_lemire(exponent, w + 1) != bits)) - { - return false; - } + digits.multiply_add(powers_of_ten[static_cast(chunk_digits)], chunk); } - bits |= negative ? (std::uint64_t{1} << 63u) : 0u; - static_assert(sizeof(double) == sizeof(std::uint64_t), "double must have 64 bits"); - std::memcpy(&out, &bits, sizeof(out)); - return true; + if (p != last) + { + point += parse_float_exponent(p + 1, last); + } + const std::int64_t s = point - count; // v = D * 10^s + + // the midpoint above the lower candidate + constexpr int mantissa_bits = Format::mantissa_bits(); + const std::uint64_t exponent_field = lower >> mantissa_bits; + std::uint64_t mantissa = lower & ((std::uint64_t{1} << mantissa_bits) - 1); + std::int64_t e = 1 + Format::minimum_exponent() - mantissa_bits; // of the smallest subnormal number + if (exponent_field != 0) + { + mantissa |= std::uint64_t{1} << mantissa_bits; + e += static_cast(exponent_field) - 1; + } + float_bigint midpoint((2 * mantissa) + 1); + const std::int64_t midpoint_exponent = e - 1; // m = midpoint * 2^midpoint_exponent + + // compare D * 5^s * 2^s with midpoint * 2^midpoint_exponent + if (s >= 0) + { + digits.multiply_power_of_five(s); + } + else + { + midpoint.multiply_power_of_five(-s); + } + const std::int64_t shift = s - midpoint_exponent; + if (shift >= 0) + { + digits.shift_left(shift); + } + else + { + midpoint.shift_left(-shift); + } + const int order = digits.compare(midpoint); + const bool round_up = order > 0 || (order == 0 && (truncated || (mantissa & 1u) != 0)); + return lower + (round_up ? 1u : 0u); } -/// Eisel-Lemire is only implemented for `double` -template -bool parse_float_eisel_lemire(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept +/// the double with the IEEE-754 bits @a bits +inline void float_from_bits(std::uint64_t bits, double& value) noexcept { - return false; + static_assert(sizeof(double) == sizeof(std::uint64_t), "double must have 64 bits"); + std::memcpy(&value, &bits, sizeof(value)); +} + +/// the float with the IEEE-754 bits @a bits (the lower 32) +inline void float_from_bits(std::uint64_t bits, float& value) noexcept +{ + static_assert(sizeof(float) == sizeof(std::uint32_t), "float must have 32 bits"); + const auto bits32 = static_cast(bits); + std::memcpy(&value, &bits32, sizeof(value)); +} + +/// the powers of ten that are exact in binary64 (up to 10^22) +inline double exact_power_of_ten(std::int64_t n, double /*tag*/) noexcept +{ + static const std::array powers = + { + { + 1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11, + 1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22 + } + }; + return powers[static_cast(n)]; +} + +/// the powers of ten that are exact in binary32 (up to 10^10) +inline float exact_power_of_ten(std::int64_t n, float /*tag*/) noexcept +{ + static const std::array powers = + { + {1e0f, 1e1f, 1e2f, 1e3f, 1e4f, 1e5f, 1e6f, 1e7f, 1e8f, 1e9f, 1e10f} + }; + return powers[static_cast(n)]; } /*! -@brief check whether Clinger's fast path can still succeed for a float token +@brief the binary32/binary64 value of a significand that was not truncated -parse_float_fast() needs a significand below 2^53. A mantissa with 17 or -more significant digits is at least 10^16 and therefore always exceeds it, -so calling the fast path would walk the token one extra time only to -decline before strtod has to run anyway. +The result is (-1)^negative * w * 10^exponent, correctly rounded (ties to +even): with Clinger's fast path where w and 10^|exponent| are exact, so that a +single floating-point operation rounds (only where intermediate results are not +kept in extended precision, see FLT_EVAL_METHOD), and with eisel_lemire() +otherwise. A value too large for the type becomes ±infinity, a value too small +±0. -Significant digits are the mantissa's digits from the first nonzero one on; -the sign, the decimal point, leading zeros, and the exponent do not count. -The answer is derived from indices - the digits are not scanned again - so -this stays off the hot path of the number scanners. +This is the core of the conversion that other parsers of JSON text share: they +can split a token themselves and still get the lexer's result. It is always +inlined, so that their hot loops keep the whole conversion inline. -@param[in] token the validated number token ('.' as decimal point) -@param[in] decimal_point_position index of the '.' in @a token, or - std::string::npos if there is none -@param[in] mantissa_end offset just past the last mantissa byte -@return false if parse_float_fast() is guaranteed to decline +@param[in] s the significand, with s.truncated == false */ -inline bool mantissa_fits_clinger(const char* token, std::size_t decimal_point_position, std::size_t mantissa_end) noexcept +template +JSON_HEDLEY_ALWAYS_INLINE FloatType decimal_to_float(const float_significand& s) noexcept { - // 10^16 already exceeds 2^53, so 17 digits can never fit - constexpr std::size_t limit = 17; + using result_type = native_float_t; + using format = ieee_binary_format::digits>; + JSON_ASSERT(!s.truncated); - const std::size_t neg = (token[0] == '-') ? 1u : 0u; - const std::size_t has_dot = (decimal_point_position != std::string::npos) ? 1u : 0u; - // the JSON grammar restricts the integer part to "0" or [1-9][0-9]*, so - // a leading zero can only be a lone "0", which is not significant - const std::size_t lead_zero = (token[neg] == '0') ? 1u : 0u; - JSON_ASSERT(mantissa_end >= neg + has_dot + lead_zero); - std::size_t digits = mantissa_end - neg - has_dot - lead_zero; - - if (JSON_HEDLEY_LIKELY(digits < limit)) +#if !defined(FLT_EVAL_METHOD) || FLT_EVAL_METHOD == 0 + if (s.exponent >= -format::max_exponent_fast_path() && s.exponent <= format::max_exponent_fast_path() + && s.w <= format::max_mantissa_fast_path()) { - return true; - } - - // Only a number below 1 can carry further insignificant zeros, and only - // while the count stays at the limit does removing them change the - // answer - so this loop is skipped for all but a few tokens. The - // fraction is located through decimal_point_position rather than by - // searching '.'. - if (lead_zero != 0) - { - JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit - for (std::size_t i = decimal_point_position + 1; - digits >= limit && i < mantissa_end && token[i] == '0'; ++i) + auto value = static_cast(s.w); + if (s.exponent < 0) { - --digits; + value /= exact_power_of_ten(-s.exponent, result_type{}); } + else + { + value *= exact_power_of_ten(s.exponent, result_type{}); + } + const FloatType result = s.negative ? -value : value; + return result; + } +#endif + + result_type value{}; + float_from_bits(eisel_lemire(s.exponent, s.w) | (s.negative ? (std::uint64_t{1} << format::sign_bit()) : 0u), value); + const FloatType result = value; + return result; +} + +/*! +@brief convert a validated number token to the nearest binary32/binary64 value + +The conversion is correctly rounded (ties to even) and independent of the +locale and of the C and C++ libraries: +1. parse_float_significand() splits the token into w * 10^q. +2. If no digits were dropped, decimal_to_float() rounds w * 10^q (Clinger's + fast path or Eisel-Lemire). +3. Otherwise, the value lies in [w, w + 1) * 10^q: if eisel_lemire() rounds + both ends to the same value, so does the token (in all but rare cases). +4. Otherwise, digit_comparison() compares the token exactly with the midpoint + between the two candidates. +A value too large for the type becomes ±infinity (the parser reports +out_of_range.406), a value too small ±0. + +@param[in] first pointer to the first character of the token +@param[in] last pointer past the last character +@param[in] decimal_point_position index of the '.' in the token, or + std::string::npos if there is none +@param[in] mantissa_end index of the 'e'/'E', or the token length +*/ +template +FloatType parse_float_native(const char* first, const char* last, + std::size_t decimal_point_position, std::size_t mantissa_end) noexcept +{ + using result_type = native_float_t; + using format = ieee_binary_format::digits>; + static_assert(std::numeric_limits::digits == std::numeric_limits::digits, "unexpected float format"); + + const float_significand s = parse_float_significand(first, last, decimal_point_position, mantissa_end); + if (JSON_HEDLEY_LIKELY(!s.truncated)) + { + return decimal_to_float(s); } - return digits < limit; + std::uint64_t bits = eisel_lemire(s.exponent, s.w); + if (JSON_HEDLEY_UNLIKELY(bits != eisel_lemire(s.exponent, s.w + 1))) + { + bits = digit_comparison(first, last, bits); + } + result_type value{}; + float_from_bits(bits | (s.negative ? (std::uint64_t{1} << format::sign_bit()) : 0u), value); + const FloatType result = value; + return result; +} + +/*! +@brief parse a float with std::from_chars when available + +Only used for the formats parse_float_native() does not convert (long double +formats other than binary64). std::from_chars is locale-independent and +correctly rounded. It is used only when __cpp_lib_to_chars indicates full +floating-point support and only when it consumes the entire token ([first, +last)). An under-/overflow (result_out_of_range) also declines, so the +caller's strtold fallback supplies the well-defined ±inf/0 result the parser +expects (side-stepping the P4168 divergence between implementations). + +@return true if the value was parsed exactly and fully; false to fall back +*/ +template +bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept +{ + // JSON_HAS_CPP_17 must gate the use as well as the include above: + // some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even + // in C++14 mode, where is not included. +#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars) + const auto result = std::from_chars(first, last, out); + return result.ec == std::errc() && result.ptr == last; +#else + static_cast(first); + static_cast(last); + static_cast(out); + return false; +#endif +} + +/// binary32 and binary64: the library's own conversion, which always succeeds +template +bool convert_float_fast(const char* first, const char* last, std::size_t decimal_point_position, + std::size_t mantissa_end, FloatType& value, std::true_type /*native*/) noexcept +{ + value = parse_float_native(first, last, decimal_point_position, mantissa_end); + return true; +} + +/// other formats (long double on x87, binary128, double-double): std::from_chars, if available +template +bool convert_float_fast(const char* first, const char* last, std::size_t /*decimal_point_position*/, + std::size_t /*mantissa_end*/, FloatType& value, std::false_type /*native*/) noexcept +{ + return parse_float_from_chars(first, last, value); } /*! @brief convert a validated float token without the C library, if possible -Tries std::from_chars (when available), Clinger's exact fast path (double -only, skipped when it cannot succeed), and the Eisel-Lemire algorithm (double -only). +float, double, and long double where it is binary64 are always converted, by +parse_float_native(). Other long double formats are converted with +std::from_chars where the standard library supports it. @param[in] first pointer to the first character of the token @param[in] last pointer past the last character @@ -9545,19 +9833,8 @@ template bool convert_float_fast(const char* first, const char* last, std::size_t decimal_point_position, std::size_t mantissa_end, FloatType& value) noexcept { - if (parse_float_from_chars(first, last, value)) - { - return true; - } - // Skipping a fast path that cannot succeed is lossless and saves a full - // extra pass over the token's bytes, which otherwise shows up on - // high-precision inputs such as canada.json - if (mantissa_fits_clinger(first, decimal_point_position, mantissa_end) - && parse_float_fast(first, last, value)) - { - return true; - } - return parse_float_eisel_lemire(first, last, value); + return convert_float_fast(first, last, decimal_point_position, mantissa_end, value, + std::integral_constant::value> {}); } /// std::strtof, std::strtod, or std::strtold, chosen by the type of @a f @@ -9592,6 +9869,11 @@ inline char get_decimal_point() noexcept /*! @brief convert a validated float token with strtof/strtod/strtold +Only used for what convert_float_fast() does not convert: long double formats +other than binary64 where std::from_chars is unavailable or reports an under- +or overflow, and floating-point types that are not IEEE-754 (see +has_native_float_format). + These functions expect the decimal point of the *current* locale, so it is looked up right before the conversion instead of once when the lexer is constructed: a locale change in between (by a parser callback, a SAX @@ -9652,6 +9934,34 @@ void convert_float_locale_aware(StringType& token, std::size_t decimal_point_pos } } +/*! +@brief convert a validated float token like the lexer does + +For parsers of JSON text other than the lexer, which converts its own token +buffer in place. float, double, and long double where it is binary64 are +converted without allocation and independent of the locale; only other long +double formats that std::from_chars does not support need a copy of the token +for convert_float_locale_aware(). + +@param[in] first pointer to the first character of the token +@param[in] last pointer past the last character +@param[in] decimal_point_position index of the '.' in the token, or + std::string::npos if there is none +@param[in] mantissa_end index of the 'e'/'E', or the token length +@return the value, ±infinity if it overflows +*/ +template +FloatType convert_float(const char* first, const char* last, std::size_t decimal_point_position, std::size_t mantissa_end) +{ + FloatType value{}; + if (!convert_float_fast(first, last, decimal_point_position, mantissa_end, value)) + { + std::string token(first, last); + convert_float_locale_aware(token, decimal_point_position, value); + } + return value; +} + } // namespace detail NLOHMANN_JSON_NAMESPACE_END @@ -11025,9 +11335,11 @@ class lexer : public lexer_base token_type::parse_error otherwise @note The scanner is independent of the current locale: token_buffer - always holds `.`. Only the std::strtod fallback of convert_number() - depends on the locale, and it looks up the decimal point right - before converting (see detail::convert_float_locale_aware()). + always holds `.`. The conversion of float and double does not use + the locale either. Only the std::strtold fallback of + convert_number() for long double formats other than binary64 + depends on it, and it looks up the decimal point right before + converting (see detail::convert_float_locale_aware()). */ token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated. { @@ -11040,7 +11352,7 @@ class lexer : public lexer_base // offset just past the last mantissa byte in token_buffer (i.e. the // index of 'e'/'E', or the whole token when there is no exponent). - // convert_number() uses it to count significant digits; npos means + // convert_number() uses it to split the token; npos means // "not seen an exponent yet" and is resolved at scan_number_done std::size_t mantissa_end = std::string::npos; @@ -11370,8 +11682,8 @@ scan_number_done: @param[in] mantissa_end offset just past the last mantissa byte in token_buffer (the index of 'e'/'E', or token_buffer.size() when there is no exponent); - used to skip Clinger's fast path when it cannot - possibly succeed - see detail::mantissa_fits_clinger() + with decimal_point_position, it locates the parts + of a float token without scanning it again */ token_type convert_number(token_type number_type, std::size_t mantissa_end) { @@ -11440,10 +11752,11 @@ scan_number_done: } // this code is reached if we parse a floating-point number or if an - // integer conversion above overflowed. Prefer std::from_chars - // (Eisel-Lemire, locale-independent, correctly rounded) when available; - // otherwise the exact Clinger fast path (double only); otherwise the - // locale-aware strtof/strtod/strtold. + // integer conversion above overflowed. float and double (and long + // double where it is binary64) are converted by the library itself, + // correctly rounded and independent of the locale; other long double + // formats use std::from_chars when available, otherwise the + // locale-aware strtold. if (convert_float_fast(num_begin, num_end, decimal_point_position, mantissa_end, value_float)) { return token_type::value_float; diff --git a/tests/src/float_hard_cases.hpp b/tests/src/float_hard_cases.hpp new file mode 100644 index 000000000..e579bccce --- /dev/null +++ b/tests/src/float_hard_cases.hpp @@ -0,0 +1,599 @@ +// __ _____ _____ _____ +// __| | __| | | | JSON for Modern C++ (supporting code) +// | | |__ | | | | | | version 3.12.0 +// |_____|_____|_____|_|___| https://github.com/nlohmann/json +// +// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann +// SPDX-License-Identifier: MIT + +#pragma once + +#include // array +#include // uint32_t, uint64_t + +// Number tokens that are hard to round correctly, with the IEEE-754 binary64 +// and binary32 bits of their correctly rounded values (ties to even; infinity +// for an overflow, a signed zero for an underflow). +// +// For doubles and floats around 0, the smallest normal number, 1, 2^24, 2^53, +// 0.1, and the largest finite number, and for random ones, the exact midpoint +// m to the next number gives: m, m with one unit more and less in the last +// digit, m with "01" and "0...01" appended, m with trailing zeros, and m cut +// after 17 to 30 digits (rounded down and up, so that the rounding is decided +// after the 19th digit), in fixed and exponent notation, 30% of them negative. +// Tokens longer than 80 characters are left out, except for four of 700 digits +// and more. Zeros, underflow, overflow, huge exponents, and integers beyond 64 +// bits complete the set. Of the 508 tokens, 134 (as double) and 150 (as +// float) need the exact comparison with the midpoint (detail::digit_comparison()). +// +// The expected bits were computed with exact rational arithmetic in Python +// (fractions.Fraction) and cross-checked with Python's float(); strtod_l and +// strtof_l of Apple's libc and of glibc agree. Generated by +// compact_hard_cases.py 5 (with hard_cases.py), see the pull request that +// added this file. + +namespace float_hard_cases +{ + +struct hard_case +{ + const char* token; + std::uint64_t bits64; + std::uint32_t bits32; +}; + +inline const std::array& cases() +{ + static const std::array table = + { + { + {"-2.4703282292062327e-324", 0x8000000000000000u, 0x80000000u}, + {"24703282292062328e-340", 0x0000000000000001u, 0x00000000u}, + {"247032822920623272e-341", 0x0000000000000000u, 0x00000000u}, + {"-0.2470328229206232721e-323", 0x8000000000000001u, 0x80000000u}, + {"-0.24703282292062327208e-323", 0x8000000000000000u, 0x80000000u}, + {"-2.4703282292062327209e-324", 0x8000000000000001u, 0x80000000u}, + {"2.47032822920623272088e-324", 0x0000000000000000u, 0x00000000u}, + {"247032822920623272089e-344", 0x0000000000000001u, 0x00000000u}, + {"-247032822920623272088284396434e-353", 0x8000000000000000u, 0x80000000u}, + {"0.247032822920623272088284396435e-323", 0x0000000000000001u, 0x00000000u}, + {"-74109846876186981e-340", 0x8000000000000001u, 0x80000000u}, + {"0.74109846876186982e-323", 0x0000000000000002u, 0x00000000u}, + {"-0.7410984687618698162e-323", 0x8000000000000001u, 0x80000000u}, + {"-7.410984687618698163e-324", 0x8000000000000002u, 0x80000000u}, + {"7.4109846876186981626e-324", 0x0000000000000001u, 0x00000000u}, + {"-74109846876186981627e-343", 0x8000000000000002u, 0x80000000u}, + {"-741098468761869816264e-344", 0x8000000000000001u, 0x80000000u}, + {"0.741098468761869816265e-323", 0x0000000000000002u, 0x00000000u}, + {"0.741098468761869816264853189302e-323", 0x0000000000000001u, 0x00000000u}, + {"-7.41098468761869816264853189303e-324", 0x8000000000000002u, 0x80000000u}, + {"0.22250738585072006e-307", 0x000FFFFFFFFFFFFEu, 0x00000000u}, + {"2.2250738585072007e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u}, + {"2.225073858507200641e-308", 0x000FFFFFFFFFFFFEu, 0x00000000u}, + {"-2225073858507200642e-326", 0x800FFFFFFFFFFFFFu, 0x80000000u}, + {"22250738585072006419e-327", 0x000FFFFFFFFFFFFEu, 0x00000000u}, + {"0.2225073858507200642e-307", 0x000FFFFFFFFFFFFFu, 0x00000000u}, + {"0.222507385850720064199e-307", 0x000FFFFFFFFFFFFEu, 0x00000000u}, + {"2.225073858507200642e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u}, + {"-2.22507385850720064199176395546e-308", 0x800FFFFFFFFFFFFEu, 0x80000000u}, + {"222507385850720064199176395547e-337", 0x000FFFFFFFFFFFFFu, 0x00000000u}, + {"-2.2250738585072011e-308", 0x800FFFFFFFFFFFFFu, 0x80000000u}, + {"-22250738585072012e-324", 0x8010000000000000u, 0x80000000u}, + {"-2225073858507201136e-326", 0x800FFFFFFFFFFFFFu, 0x80000000u}, + {"0.2225073858507201137e-307", 0x0010000000000000u, 0x00000000u}, + {"0.2225073858507201136e-307", 0x000FFFFFFFFFFFFFu, 0x00000000u}, + {"-2.2250738585072011361e-308", 0x8010000000000000u, 0x80000000u}, + {"2.22507385850720113605e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u}, + {"222507385850720113606e-328", 0x0010000000000000u, 0x00000000u}, + {"22250738585072011360574097967e-336", 0x000FFFFFFFFFFFFFu, 0x00000000u}, + {"0.222507385850720113605740979671e-307", 0x0010000000000000u, 0x00000000u}, + {"22250738585072016e-324", 0x0010000000000000u, 0x00000000u}, + {"0.22250738585072017e-307", 0x0010000000000001u, 0x00000000u}, + {"0.222507385850720163e-307", 0x0010000000000000u, 0x00000000u}, + {"2.225073858507201631e-308", 0x0010000000000001u, 0x00000000u}, + {"-2.2250738585072016301e-308", 0x8010000000000000u, 0x80000000u}, + {"22250738585072016302e-327", 0x0010000000000001u, 0x00000000u}, + {"-222507385850720163012e-328", 0x8010000000000000u, 0x80000000u}, + {"0.222507385850720163013e-307", 0x0010000000000001u, 0x00000000u}, + {"0.222507385850720163012305563795e-307", 0x0010000000000000u, 0x00000000u}, + {"-2.22507385850720163012305563796e-308", 0x8010000000000001u, 0x80000000u}, + {"0.17976931348623156E+309", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u}, + {"1.7976931348623157e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u}, + {"1.797693134862315608e308", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u}, + {"-1797693134862315609e290", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u}, + {"-17976931348623156083e289", 0xFFEFFFFFFFFFFFFEu, 0xFF800000u}, + {"-0.17976931348623156084E+309", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u}, + {"0.179769313486231560835E+309", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u}, + {"-1.79769313486231560836e308", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u}, + {"1.79769313486231560835325876058e308", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u}, + {"179769313486231560835325876059e279", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u}, + {"1.7976931348623158e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u}, + {"17976931348623159e292", 0x7FF0000000000000u, 0x7F800000u}, + {"1797693134862315807e290", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u}, + {"0.1797693134862315808E+309", 0x7FF0000000000000u, 0x7F800000u}, + {"0.17976931348623158079E+309", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u}, + {"-1.797693134862315808e308", 0xFFF0000000000000u, 0xFF800000u}, + {"1.79769313486231580793e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u}, + {"179769313486231580794e288", 0x7FF0000000000000u, 0x7F800000u}, + {"179769313486231580793728971405e279", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u}, + {"-0.179769313486231580793728971406E+309", 0xFFF0000000000000u, 0xFF800000u}, + {"100000000000000011102230246251565404236316680908203125e-53", 0x3FF0000000000000u, 0x3F800000u}, + {"-1.00000000000000011102230246251565404236316680908203126", 0xBFF0000000000001u, 0xBF800000u}, + {"1.00000000000000011102230246251565404236316680908203124e0", 0x3FF0000000000000u, 0x3F800000u}, + {"10000000000000001110223024625156540423631668090820312501e-55", 0x3FF0000000000001u, 0x3F800000u}, + {"1.00000000000000011102230246251565404236316680908203125000000000000000000001", 0x3FF0000000000001u, 0x3F800000u}, + {"10000000000000001e-16", 0x3FF0000000000000u, 0x3F800000u}, + {"1.0000000000000002", 0x3FF0000000000001u, 0x3F800000u}, + {"1.000000000000000111", 0x3FF0000000000000u, 0x3F800000u}, + {"1.000000000000000112e0", 0x3FF0000000000001u, 0x3F800000u}, + {"1.000000000000000111e0", 0x3FF0000000000000u, 0x3F800000u}, + {"-10000000000000001111e-19", 0xBFF0000000000001u, 0xBF800000u}, + {"-100000000000000011102e-20", 0xBFF0000000000000u, 0xBF800000u}, + {"-1.00000000000000011103", 0xBFF0000000000001u, 0xBF800000u}, + {"1.00000000000000011102230246251", 0x3FF0000000000000u, 0x3F800000u}, + {"1.00000000000000011102230246252e0", 0x3FF0000000000001u, 0x3F800000u}, + {"-0.999999999999999944488848768742172978818416595458984375", 0xBFF0000000000000u, 0xBF800000u}, + {"-9.99999999999999944488848768742172978818416595458984376e-1", 0xBFF0000000000000u, 0xBF800000u}, + {"999999999999999944488848768742172978818416595458984374e-54", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u}, + {"0.99999999999999994448884876874217297881841659545898437501", 0x3FF0000000000000u, 0x3F800000u}, + {"9.99999999999999944488848768742172978818416595458984375000000000000000000001e-1", 0x3FF0000000000000u, 0x3F800000u}, + {"-0.99999999999999994", 0xBFEFFFFFFFFFFFFFu, 0xBF800000u}, + {"9.9999999999999995e-1", 0x3FF0000000000000u, 0x3F800000u}, + {"9.999999999999999444e-1", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u}, + {"9999999999999999445e-19", 0x3FF0000000000000u, 0x3F800000u}, + {"99999999999999994448e-20", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u}, + {"-0.99999999999999994449", 0xBFF0000000000000u, 0xBF800000u}, + {"0.999999999999999944488", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u}, + {"-9.99999999999999944489e-1", 0xBFF0000000000000u, 0xBF800000u}, + {"9.99999999999999944488848768742e-1", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u}, + {"999999999999999944488848768743e-30", 0x3FF0000000000000u, 0x3F800000u}, + {"-9.007199254740993e15", 0xC340000000000000u, 0xDA000000u}, + {"9007199254740994e0", 0x4340000000000001u, 0x5A000000u}, + {"9007199254740992", 0x4340000000000000u, 0x5A000000u}, + {"9.00719925474099301e15", 0x4340000000000001u, 0x5A000000u}, + {"9007199254740993000000000000000000001e-21", 0x4340000000000001u, 0x5A000000u}, + {"-9007199254740993.000000000000000000000000000000", 0xC340000000000000u, 0xDA000000u}, + {"90071992547409915e-1", 0x4340000000000000u, 0x5A000000u}, + {"-9007199254740991.6", 0xC340000000000000u, 0xDA000000u}, + {"9.0071992547409914e15", 0x433FFFFFFFFFFFFFu, 0x5A000000u}, + {"-9007199254740991501e-3", 0xC340000000000000u, 0xDA000000u}, + {"9007199254740991.5000000000000000000001", 0x4340000000000000u, 0x5A000000u}, + {"9.0071992547409915000000000000000000000000000000e15", 0x4340000000000000u, 0x5A000000u}, + {"0.100000000000000012490009027033011079765856266021728515625", 0x3FB999999999999Au, 0x3DCCCCCDu}, + {"1.00000000000000012490009027033011079765856266021728515626e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu}, + {"100000000000000012490009027033011079765856266021728515624e-57", 0x3FB999999999999Au, 0x3DCCCCCDu}, + {"0.10000000000000001249000902703301107976585626602172851562501", 0x3FB999999999999Bu, 0x3DCCCCCDu}, + {"0.10000000000000001", 0x3FB999999999999Au, 0x3DCCCCCDu}, + {"1.0000000000000002e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu}, + {"-1.000000000000000124e-1", 0xBFB999999999999Au, 0xBDCCCCCDu}, + {"1000000000000000125e-19", 0x3FB999999999999Bu, 0x3DCCCCCDu}, + {"-10000000000000001249e-20", 0xBFB999999999999Au, 0xBDCCCCCDu}, + {"0.1000000000000000125", 0x3FB999999999999Bu, 0x3DCCCCCDu}, + {"0.10000000000000001249", 0x3FB999999999999Au, 0x3DCCCCCDu}, + {"1.00000000000000012491e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu}, + {"1.00000000000000012490009027033e-1", 0x3FB999999999999Au, 0x3DCCCCCDu}, + {"100000000000000012490009027034e-30", 0x3FB999999999999Bu, 0x3DCCCCCDu}, + {"2.45134755833537796875e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u}, + {"-245134755833537796876e-6", 0xC2EBDE5C4164D83Au, 0xD75EF2E2u}, + {"-245134755833537.796874", 0xC2EBDE5C4164D839u, 0xD75EF2E2u}, + {"2.4513475583353779687501e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u}, + {"245134755833537796875000000000000000000001e-27", 0x42EBDE5C4164D83Au, 0x575EF2E2u}, + {"245134755833537.796875000000000000000000000000000000", 0x42EBDE5C4164D83Au, 0x575EF2E2u}, + {"2.4513475583353779e14", 0x42EBDE5C4164D839u, 0x575EF2E2u}, + {"2451347558335378e-1", 0x42EBDE5C4164D83Au, 0x575EF2E2u}, + {"2451347558335377968e-4", 0x42EBDE5C4164D839u, 0x575EF2E2u}, + {"245134755833537.7969", 0x42EBDE5C4164D83Au, 0x575EF2E2u}, + {"245134755833537.79687", 0x42EBDE5C4164D839u, 0x575EF2E2u}, + {"2.4513475583353779688e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u}, + {"181510327827821147441864013671875e-23", 0x41DB0C11CB91CE38u, 0x4ED8608Eu}, + {"-1815103278.27821147441864013671876", 0xC1DB0C11CB91CE38u, 0xCED8608Eu}, + {"1.81510327827821147441864013671874e9", 0x41DB0C11CB91CE37u, 0x4ED8608Eu}, + {"18151032782782114744186401367187501e-25", 0x41DB0C11CB91CE38u, 0x4ED8608Eu}, + {"1815103278.27821147441864013671875000000000000000000001", 0x41DB0C11CB91CE38u, 0x4ED8608Eu}, + {"-1.81510327827821147441864013671875000000000000000000000000000000e9", 0xC1DB0C11CB91CE38u, 0xCED8608Eu}, + {"18151032782782114e-7", 0x41DB0C11CB91CE37u, 0x4ED8608Eu}, + {"1815103278.2782115", 0x41DB0C11CB91CE38u, 0x4ED8608Eu}, + {"1815103278.278211474", 0x41DB0C11CB91CE37u, 0x4ED8608Eu}, + {"-1.815103278278211475e9", 0xC1DB0C11CB91CE38u, 0xCED8608Eu}, + {"1.8151032782782114744e9", 0x41DB0C11CB91CE37u, 0x4ED8608Eu}, + {"-18151032782782114745e-10", 0xC1DB0C11CB91CE38u, 0xCED8608Eu}, + {"181510327827821147441e-11", 0x41DB0C11CB91CE37u, 0x4ED8608Eu}, + {"1815103278.27821147442", 0x41DB0C11CB91CE38u, 0x4ED8608Eu}, + {"1815103278.27821147441864013671", 0x41DB0C11CB91CE37u, 0x4ED8608Eu}, + {"1.81510327827821147441864013672e9", 0x41DB0C11CB91CE38u, 0x4ED8608Eu}, + {"3809325632181785344", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u}, + {"3.809325632181785345e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u}, + {"3809325632181785343e0", 0x43CA6EB8BD69FE29u, 0x5E5375C6u}, + {"3809325632181785344.01", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u}, + {"3.809325632181785344000000000000000000001e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u}, + {"3809325632181785344000000000000000000000000000000e-30", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u}, + {"3809325632181785300", 0x43CA6EB8BD69FE29u, 0x5E5375C6u}, + {"3.8093256321817854e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u}, + {"4.046966549916366943359375e12", 0x428D7210076CE2F0u, 0x546B9080u}, + {"4046966549916366943359376e-12", 0x428D7210076CE2F0u, 0x546B9080u}, + {"4046966549916.366943359374", 0x428D7210076CE2EFu, 0x546B9080u}, + {"4.04696654991636694335937501e12", 0x428D7210076CE2F0u, 0x546B9080u}, + {"4046966549916366943359375000000000000000000001e-33", 0x428D7210076CE2F0u, 0x546B9080u}, + {"-4046966549916.366943359375000000000000000000000000000000", 0xC28D7210076CE2F0u, 0xD46B9080u}, + {"4.0469665499163669e12", 0x428D7210076CE2EFu, 0x546B9080u}, + {"4046966549916367e-3", 0x428D7210076CE2F0u, 0x546B9080u}, + {"4046966549916366943e-6", 0x428D7210076CE2EFu, 0x546B9080u}, + {"4046966549916.366944", 0x428D7210076CE2F0u, 0x546B9080u}, + {"4046966549916.3669433", 0x428D7210076CE2EFu, 0x546B9080u}, + {"4.0469665499163669434e12", 0x428D7210076CE2F0u, 0x546B9080u}, + {"-4.04696654991636694335e12", 0xC28D7210076CE2EFu, 0xD46B9080u}, + {"404696654991636694336e-8", 0x428D7210076CE2F0u, 0x546B9080u}, + {"28093802557000874e154", 0x63529C3B77330BDBu, 0x7F800000u}, + {"0.28093802557000875E+171", 0x63529C3B77330BDCu, 0x7F800000u}, + {"0.2809380255700087447E+171", 0x63529C3B77330BDBu, 0x7F800000u}, + {"2.809380255700087448e170", 0x63529C3B77330BDCu, 0x7F800000u}, + {"2.8093802557000874472e170", 0x63529C3B77330BDBu, 0x7F800000u}, + {"28093802557000874473e151", 0x63529C3B77330BDCu, 0x7F800000u}, + {"280938025570008744728e150", 0x63529C3B77330BDBu, 0x7F800000u}, + {"-0.280938025570008744729E+171", 0xE3529C3B77330BDCu, 0xFF800000u}, + {"-0.280938025570008744728403667979E+171", 0xE3529C3B77330BDBu, 0xFF800000u}, + {"2.8093802557000874472840366798e170", 0x63529C3B77330BDCu, 0x7F800000u}, + {"0.39523280297734525e-154", 0x1FE0F51BF17FD374u, 0x00000000u}, + {"-3.9523280297734526e-155", 0x9FE0F51BF17FD375u, 0x80000000u}, + {"3.952328029773452547e-155", 0x1FE0F51BF17FD374u, 0x00000000u}, + {"3952328029773452548e-173", 0x1FE0F51BF17FD375u, 0x00000000u}, + {"-39523280297734525478e-174", 0x9FE0F51BF17FD374u, 0x80000000u}, + {"0.39523280297734525479e-154", 0x1FE0F51BF17FD375u, 0x00000000u}, + {"-0.395232802977345254787e-154", 0x9FE0F51BF17FD374u, 0x80000000u}, + {"-3.95232802977345254788e-155", 0x9FE0F51BF17FD375u, 0x80000000u}, + {"-3.95232802977345254787245825501e-155", 0x9FE0F51BF17FD374u, 0x80000000u}, + {"395232802977345254787245825502e-184", 0x1FE0F51BF17FD375u, 0x00000000u}, + {"-1.0790205420931879e-276", 0x86A3209CA6233255u, 0x80000000u}, + {"-1079020542093188e-291", 0x86A3209CA6233256u, 0x80000000u}, + {"-1079020542093187947e-294", 0x86A3209CA6233255u, 0x80000000u}, + {"0.1079020542093187948e-275", 0x06A3209CA6233256u, 0x00000000u}, + {"0.1079020542093187947e-275", 0x06A3209CA6233255u, 0x00000000u}, + {"1.0790205420931879471e-276", 0x06A3209CA6233256u, 0x00000000u}, + {"1.07902054209318794701e-276", 0x06A3209CA6233255u, 0x00000000u}, + {"-107902054209318794702e-296", 0x86A3209CA6233256u, 0x80000000u}, + {"107902054209318794701153285302e-305", 0x06A3209CA6233255u, 0x00000000u}, + {"-0.107902054209318794701153285303e-275", 0x86A3209CA6233256u, 0x80000000u}, + {"58530471071351308e-228", 0x1413B446E6A16A3Bu, 0x00000000u}, + {"-0.58530471071351309e-211", 0x9413B446E6A16A3Cu, 0x80000000u}, + {"0.5853047107135130893e-211", 0x1413B446E6A16A3Bu, 0x00000000u}, + {"-5.853047107135130894e-212", 0x9413B446E6A16A3Cu, 0x80000000u}, + {"5.853047107135130893e-212", 0x1413B446E6A16A3Bu, 0x00000000u}, + {"58530471071351308931e-231", 0x1413B446E6A16A3Cu, 0x00000000u}, + {"-585304710713513089304e-232", 0x9413B446E6A16A3Bu, 0x80000000u}, + {"0.585304710713513089305e-211", 0x1413B446E6A16A3Cu, 0x00000000u}, + {"-0.585304710713513089304248824438e-211", 0x9413B446E6A16A3Bu, 0x80000000u}, + {"5.85304710713513089304248824439e-212", 0x1413B446E6A16A3Cu, 0x00000000u}, + {"0.19334214893983531e-78", 0x2F96ECBF1CFB10F6u, 0x00000000u}, + {"1.9334214893983532e-79", 0x2F96ECBF1CFB10F7u, 0x00000000u}, + {"1.933421489398353102e-79", 0x2F96ECBF1CFB10F6u, 0x00000000u}, + {"1933421489398353103e-97", 0x2F96ECBF1CFB10F7u, 0x00000000u}, + {"19334214893983531023e-98", 0x2F96ECBF1CFB10F6u, 0x00000000u}, + {"0.19334214893983531024e-78", 0x2F96ECBF1CFB10F7u, 0x00000000u}, + {"0.193342148939835310231e-78", 0x2F96ECBF1CFB10F6u, 0x00000000u}, + {"1.93342148939835310232e-79", 0x2F96ECBF1CFB10F7u, 0x00000000u}, + {"-1.93342148939835310231359014704e-79", 0xAF96ECBF1CFB10F6u, 0x80000000u}, + {"193342148939835310231359014705e-108", 0x2F96ECBF1CFB10F7u, 0x00000000u}, + {"2.9873358928024455e227", 0x6F2938807814E8A2u, 0x7F800000u}, + {"29873358928024456e211", 0x6F2938807814E8A3u, 0x7F800000u}, + {"298733589280244551e210", 0x6F2938807814E8A2u, 0x7F800000u}, + {"0.2987335892802445511E+228", 0x6F2938807814E8A3u, 0x7F800000u}, + {"0.29873358928024455109E+228", 0x6F2938807814E8A2u, 0x7F800000u}, + {"2.987335892802445511e227", 0x6F2938807814E8A3u, 0x7F800000u}, + {"-2.98733589280244551098e227", 0xEF2938807814E8A2u, 0xFF800000u}, + {"-298733589280244551099e207", 0xEF2938807814E8A3u, 0xFF800000u}, + {"298733589280244551098081559931e198", 0x6F2938807814E8A2u, 0x7F800000u}, + {"0.298733589280244551098081559932E+228", 0x6F2938807814E8A3u, 0x7F800000u}, + {"7.0064923216240853e-46", 0x3690000000000000u, 0x00000000u}, + {"70064923216240854e-62", 0x3690000000000000u, 0x00000001u}, + {"7006492321624085354e-64", 0x3690000000000000u, 0x00000000u}, + {"-0.7006492321624085355e-45", 0xB690000000000000u, 0x80000001u}, + {"0.70064923216240853546e-45", 0x3690000000000000u, 0x00000000u}, + {"7.0064923216240853547e-46", 0x3690000000000000u, 0x00000001u}, + {"7.00649232162408535461e-46", 0x3690000000000000u, 0x00000000u}, + {"-700649232162408535462e-66", 0xB690000000000000u, 0x80000001u}, + {"-700649232162408535461864791644e-75", 0xB690000000000000u, 0x80000000u}, + {"0.700649232162408535461864791645e-45", 0x3690000000000000u, 0x00000001u}, + {"21019476964872256e-61", 0x36A8000000000000u, 0x00000001u}, + {"0.21019476964872257e-44", 0x36A8000000000000u, 0x00000002u}, + {"-0.2101947696487225606e-44", 0xB6A8000000000000u, 0x80000001u}, + {"-2.101947696487225607e-45", 0xB6A8000000000000u, 0x80000002u}, + {"2.1019476964872256063e-45", 0x36A8000000000000u, 0x00000001u}, + {"21019476964872256064e-64", 0x36A8000000000000u, 0x00000002u}, + {"-210194769648722560638e-65", 0xB6A8000000000000u, 0x80000001u}, + {"-0.210194769648722560639e-44", 0xB6A8000000000000u, 0x80000002u}, + {"-0.210194769648722560638559437493e-44", 0xB6A8000000000000u, 0x80000001u}, + {"2.10194769648722560638559437494e-45", 0x36A8000000000000u, 0x00000002u}, + {"0.11754941406275178e-37", 0x380FFFFFA0000000u, 0x007FFFFEu}, + {"1.1754941406275179e-38", 0x380FFFFFA0000000u, 0x007FFFFFu}, + {"-1.175494140627517859e-38", 0xB80FFFFFA0000000u, 0x807FFFFEu}, + {"117549414062751786e-55", 0x380FFFFFA0000000u, 0x007FFFFFu}, + {"-11754941406275178592e-57", 0xB80FFFFFA0000000u, 0x807FFFFEu}, + {"0.11754941406275178593e-37", 0x380FFFFFA0000000u, 0x007FFFFFu}, + {"0.117549414062751785924e-37", 0x380FFFFFA0000000u, 0x007FFFFEu}, + {"1.17549414062751785925e-38", 0x380FFFFFA0000000u, 0x007FFFFFu}, + {"-1.17549414062751785924617589866e-38", 0xB80FFFFFA0000000u, 0x807FFFFEu}, + {"117549414062751785924617589867e-67", 0x380FFFFFA0000000u, 0x007FFFFFu}, + {"1.1754942807573642e-38", 0x380FFFFFDFFFFFFFu, 0x007FFFFFu}, + {"11754942807573643e-54", 0x380FFFFFE0000000u, 0x00800000u}, + {"1175494280757364291e-56", 0x380FFFFFE0000000u, 0x007FFFFFu}, + {"0.1175494280757364292e-37", 0x380FFFFFE0000000u, 0x00800000u}, + {"0.11754942807573642917e-37", 0x380FFFFFE0000000u, 0x007FFFFFu}, + {"-1.1754942807573642918e-38", 0xB80FFFFFE0000000u, 0x80800000u}, + {"-1.17549428075736429172e-38", 0xB80FFFFFE0000000u, 0x807FFFFFu}, + {"-117549428075736429173e-58", 0xB80FFFFFE0000000u, 0x80800000u}, + {"117549428075736429172788299103e-67", 0x380FFFFFE0000000u, 0x007FFFFFu}, + {"0.117549428075736429172788299104e-37", 0x380FFFFFE0000000u, 0x00800000u}, + {"11754944208872107e-54", 0x3810000010000000u, 0x00800000u}, + {"-0.11754944208872108e-37", 0xB810000010000000u, 0x80800001u}, + {"0.1175494420887210724e-37", 0x3810000010000000u, 0x00800000u}, + {"1.175494420887210725e-38", 0x3810000010000000u, 0x00800001u}, + {"1.1754944208872107242e-38", 0x3810000010000000u, 0x00800000u}, + {"-11754944208872107243e-57", 0xB810000010000000u, 0x80800001u}, + {"-11754944208872107242e-57", 0xB810000010000000u, 0x80800000u}, + {"0.117549442088721072421e-37", 0x3810000010000000u, 0x00800001u}, + {"0.11754944208872107242095900834e-37", 0x3810000010000000u, 0x00800000u}, + {"-1.17549442088721072420959008341e-38", 0xB810000010000000u, 0x80800001u}, + {"340282336497324057985868971510891282432", 0x47EFFFFFD0000000u, 0x7F7FFFFEu}, + {"-3.40282336497324057985868971510891282433e38", 0xC7EFFFFFD0000000u, 0xFF7FFFFFu}, + {"340282336497324057985868971510891282431e0", 0x47EFFFFFD0000000u, 0x7F7FFFFEu}, + {"340282336497324057985868971510891282432.01", 0x47EFFFFFD0000000u, 0x7F7FFFFFu}, + {"3.40282336497324057985868971510891282432000000000000000000001e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu}, + {"340282336497324057985868971510891282432000000000000000000000000000000e-30", 0x47EFFFFFD0000000u, 0x7F7FFFFEu}, + {"340282336497324050000000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFEu}, + {"3.4028233649732406e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu}, + {"3.402823364973240579e38", 0x47EFFFFFD0000000u, 0x7F7FFFFEu}, + {"340282336497324058e21", 0x47EFFFFFD0000000u, 0x7F7FFFFFu}, + {"34028233649732405798e19", 0x47EFFFFFD0000000u, 0x7F7FFFFEu}, + {"340282336497324057990000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFFu}, + {"340282336497324057985000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFEu}, + {"3.40282336497324057986e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu}, + {"3.4028233649732405798586897151e38", 0x47EFFFFFD0000000u, 0x7F7FFFFEu}, + {"340282336497324057985868971511e9", 0x47EFFFFFD0000000u, 0x7F7FFFFFu}, + {"3.40282356779733661637539395458142568448e38", 0x47EFFFFFF0000000u, 0x7F800000u}, + {"-340282356779733661637539395458142568449e0", 0xC7EFFFFFF0000000u, 0xFF800000u}, + {"340282356779733661637539395458142568447", 0x47EFFFFFF0000000u, 0x7F7FFFFFu}, + {"3.4028235677973366163753939545814256844801e38", 0x47EFFFFFF0000000u, 0x7F800000u}, + {"340282356779733661637539395458142568448000000000000000000001e-21", 0x47EFFFFFF0000000u, 0x7F800000u}, + {"340282356779733661637539395458142568448.000000000000000000000000000000", 0x47EFFFFFF0000000u, 0x7F800000u}, + {"3.4028235677973366e38", 0x47EFFFFFF0000000u, 0x7F7FFFFFu}, + {"-34028235677973367e22", 0xC7EFFFFFF0000000u, 0xFF800000u}, + {"3402823567797336616e20", 0x47EFFFFFF0000000u, 0x7F7FFFFFu}, + {"340282356779733661700000000000000000000", 0x47EFFFFFF0000000u, 0x7F800000u}, + {"340282356779733661630000000000000000000", 0x47EFFFFFF0000000u, 0x7F7FFFFFu}, + {"3.4028235677973366164e38", 0x47EFFFFFF0000000u, 0x7F800000u}, + {"3.40282356779733661637e38", 0x47EFFFFFF0000000u, 0x7F7FFFFFu}, + {"340282356779733661638e18", 0x47EFFFFFF0000000u, 0x7F800000u}, + {"-340282356779733661637539395458e9", 0xC7EFFFFFF0000000u, 0xFF7FFFFFu}, + {"340282356779733661637539395459000000000", 0x47EFFFFFF0000000u, 0x7F800000u}, + {"-1000000059604644775390625e-24", 0xBFF0000010000000u, 0xBF800000u}, + {"-1.000000059604644775390626", 0xBFF0000010000000u, 0xBF800001u}, + {"-1.000000059604644775390624e0", 0xBFF0000010000000u, 0xBF800000u}, + {"100000005960464477539062501e-26", 0x3FF0000010000000u, 0x3F800001u}, + {"1.000000059604644775390625000000000000000000001", 0x3FF0000010000000u, 0x3F800001u}, + {"1.000000059604644775390625000000000000000000000000000000e0", 0x3FF0000010000000u, 0x3F800000u}, + {"-10000000596046447e-16", 0xBFF0000010000000u, 0xBF800000u}, + {"1.0000000596046448", 0x3FF0000010000000u, 0x3F800001u}, + {"1.000000059604644775", 0x3FF0000010000000u, 0x3F800000u}, + {"-1.000000059604644776e0", 0xBFF0000010000000u, 0xBF800001u}, + {"-1.0000000596046447753e0", 0xBFF0000010000000u, 0xBF800000u}, + {"10000000596046447754e-19", 0x3FF0000010000000u, 0x3F800001u}, + {"100000005960464477539e-20", 0x3FF0000010000000u, 0x3F800000u}, + {"-1.0000000596046447754", 0xBFF0000010000000u, 0xBF800001u}, + {"0.9999999701976776123046875", 0x3FEFFFFFF0000000u, 0x3F800000u}, + {"9.999999701976776123046876e-1", 0x3FEFFFFFF0000000u, 0x3F800000u}, + {"9999999701976776123046874e-25", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu}, + {"0.999999970197677612304687501", 0x3FEFFFFFF0000000u, 0x3F800000u}, + {"-9.999999701976776123046875000000000000000000001e-1", 0xBFEFFFFFF0000000u, 0xBF800000u}, + {"-9999999701976776123046875000000000000000000000000000000e-55", 0xBFEFFFFFF0000000u, 0xBF800000u}, + {"0.99999997019767761", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu}, + {"-9.9999997019767762e-1", 0xBFEFFFFFF0000000u, 0xBF800000u}, + {"-9.999999701976776123e-1", 0xBFEFFFFFF0000000u, 0xBF7FFFFFu}, + {"9999999701976776124e-19", 0x3FEFFFFFF0000000u, 0x3F800000u}, + {"9999999701976776123e-19", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu}, + {"0.99999997019767761231", 0x3FEFFFFFF0000000u, 0x3F800000u}, + {"-0.999999970197677612304", 0xBFEFFFFFF0000000u, 0xBF7FFFFFu}, + {"9.99999970197677612305e-1", 0x3FEFFFFFF0000000u, 0x3F800000u}, + {"-1.6777217e7", 0xC170000010000000u, 0xCB800000u}, + {"16777218e0", 0x4170000020000000u, 0x4B800001u}, + {"16777216", 0x4170000000000000u, 0x4B800000u}, + {"-1.677721701e7", 0xC17000001028F5C3u, 0xCB800001u}, + {"-16777217000000000000000000001e-21", 0xC170000010000000u, 0xCB800001u}, + {"16777217.000000000000000000000000000000", 0x4170000010000000u, 0x4B800000u}, + {"167772155e-1", 0x416FFFFFF0000000u, 0x4B800000u}, + {"16777215.6", 0x416FFFFFF3333333u, 0x4B800000u}, + {"1.67772154e7", 0x416FFFFFECCCCCCDu, 0x4B7FFFFFu}, + {"16777215501e-3", 0x416FFFFFF0083127u, 0x4B800000u}, + {"-16777215.5000000000000000000001", 0xC16FFFFFF0000000u, 0xCB800000u}, + {"-1.67772155000000000000000000000000000000e7", 0xC16FFFFFF0000000u, 0xCB800000u}, + {"0.1000000052154064178466796875", 0x3FB99999B0000000u, 0x3DCCCCCEu}, + {"1.000000052154064178466796876e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu}, + {"-1000000052154064178466796874e-28", 0xBFB99999B0000000u, 0xBDCCCCCDu}, + {"0.100000005215406417846679687501", 0x3FB99999B0000000u, 0x3DCCCCCEu}, + {"1.000000052154064178466796875000000000000000000001e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu}, + {"-1000000052154064178466796875000000000000000000000000000000e-58", 0xBFB99999B0000000u, 0xBDCCCCCEu}, + {"-0.10000000521540641", 0xBFB99999AFFFFFFFu, 0xBDCCCCCDu}, + {"-1.0000000521540642e-1", 0xBFB99999B0000000u, 0xBDCCCCCEu}, + {"1.000000052154064178e-1", 0x3FB99999B0000000u, 0x3DCCCCCDu}, + {"-1000000052154064179e-19", 0xBFB99999B0000000u, 0xBDCCCCCEu}, + {"10000000521540641784e-20", 0x3FB99999B0000000u, 0x3DCCCCCDu}, + {"0.10000000521540641785", 0x3FB99999B0000000u, 0x3DCCCCCEu}, + {"0.100000005215406417846", 0x3FB99999B0000000u, 0x3DCCCCCDu}, + {"1.00000005215406417847e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu}, + {"5.429001220703125e3", 0x40B5350050000000u, 0x45A9A802u}, + {"-5429001220703126e-12", 0xC0B5350050000001u, 0xC5A9A803u}, + {"-5429.001220703124", 0xC0B535004FFFFFFFu, 0xC5A9A802u}, + {"5.42900122070312501e3", 0x40B5350050000000u, 0x45A9A803u}, + {"5429001220703125000000000000000000001e-33", 0x40B5350050000000u, 0x45A9A803u}, + {"-5429.001220703125000000000000000000000000000000", 0xC0B5350050000000u, 0xC5A9A802u}, + {"503719056e0", 0x41BE062490000000u, 0x4DF03124u}, + {"503719057", 0x41BE062491000000u, 0x4DF03125u}, + {"5.03719055e8", 0x41BE06248F000000u, 0x4DF03124u}, + {"50371905601e-2", 0x41BE062490028F5Cu, 0x4DF03125u}, + {"503719056.000000000000000000001", 0x41BE062490000000u, 0x4DF03125u}, + {"5.03719056000000000000000000000000000000e8", 0x41BE062490000000u, 0x4DF03124u}, + {"-92331620", 0xC196037990000000u, 0xCCB01BCCu}, + {"9.233163e7", 0x41960379B8000000u, 0x4CB01BCEu}, + {"9233161e1", 0x4196037968000000u, 0x4CB01BCBu}, + {"92331620.1", 0x4196037990666666u, 0x4CB01BCDu}, + {"9.233162000000000000000000001e7", 0x4196037990000000u, 0x4CB01BCDu}, + {"9233162000000000000000000000000000000e-29", 0x4196037990000000u, 0x4CB01BCCu}, + {"3.002458625e6", 0x4146E82D50000000u, 0x4A37416Au}, + {"3002458626e-3", 0x4146E82D5020C49Cu, 0x4A37416Bu}, + {"3002458.624", 0x4146E82D4FDF3B64u, 0x4A37416Au}, + {"-3.00245862501e6", 0xC146E82D500053E3u, 0xCA37416Bu}, + {"3002458625000000000000000000001e-24", 0x4146E82D50000000u, 0x4A37416Bu}, + {"3002458.625000000000000000000000000000000", 0x4146E82D50000000u, 0x4A37416Au}, + {"-1095485584696182596504479582065262592e1", 0xC7A07BA830000000u, 0xFD03DD42u}, + {"10954855846961825965044795820652625930", 0x47A07BA830000000u, 0x7D03DD42u}, + {"1.095485584696182596504479582065262591e37", 0x47A07BA830000000u, 0x7D03DD41u}, + {"109548558469618259650447958206526259201e-1", 0x47A07BA830000000u, 0x7D03DD42u}, + {"-10954855846961825965044795820652625920.00000000000000000001", 0xC7A07BA830000000u, 0xFD03DD42u}, + {"1.095485584696182596504479582065262592000000000000000000000000000000e37", 0x47A07BA830000000u, 0x7D03DD42u}, + {"10954855846961825e21", 0x47A07BA830000000u, 0x7D03DD41u}, + {"10954855846961826000000000000000000000", 0x47A07BA830000000u, 0x7D03DD42u}, + {"10954855846961825960000000000000000000", 0x47A07BA830000000u, 0x7D03DD41u}, + {"1.095485584696182597e37", 0x47A07BA830000000u, 0x7D03DD42u}, + {"1.0954855846961825965e37", 0x47A07BA830000000u, 0x7D03DD41u}, + {"-10954855846961825966e18", 0xC7A07BA830000000u, 0xFD03DD42u}, + {"10954855846961825965e18", 0x47A07BA830000000u, 0x7D03DD41u}, + {"10954855846961825965100000000000000000", 0x47A07BA830000000u, 0x7D03DD42u}, + {"10954855846961825965044795820600000000", 0x47A07BA830000000u, 0x7D03DD41u}, + {"-1.09548558469618259650447958207e37", 0xC7A07BA830000000u, 0xFD03DD42u}, + {"1.6449216019182103706535606608388384863861375606575165875256061553955078126e-21", 0x3B9F125A50000000u, 0x1CF892D3u}, + {"-16449216019182103706535606608388384863861375606575165875256061553955078124e-94", 0xBB9F125A50000000u, 0x9CF892D2u}, + {"-0.0000000000000000000016449216019182103", 0xBB9F125A50000000u, 0x9CF892D2u}, + {"-1.6449216019182104e-21", 0xBB9F125A50000000u, 0x9CF892D3u}, + {"-1.64492160191821037e-21", 0xBB9F125A50000000u, 0x9CF892D2u}, + {"-1644921601918210371e-39", 0xBB9F125A50000000u, 0x9CF892D3u}, + {"-16449216019182103706e-40", 0xBB9F125A50000000u, 0x9CF892D2u}, + {"0.0000000000000000000016449216019182103707", 0x3B9F125A50000000u, 0x1CF892D3u}, + {"0.00000000000000000000164492160191821037065", 0x3B9F125A50000000u, 0x1CF892D2u}, + {"1.64492160191821037066e-21", 0x3B9F125A50000000u, 0x1CF892D3u}, + {"1.64492160191821037065356066083e-21", 0x3B9F125A50000000u, 0x1CF892D2u}, + {"164492160191821037065356066084e-50", 0x3B9F125A50000000u, 0x1CF892D3u}, + {"6.565061509609222412109375e-1", 0x3FE5021930000000u, 0x3F2810CAu}, + {"6565061509609222412109376e-25", 0x3FE5021930000000u, 0x3F2810CAu}, + {"0.6565061509609222412109374", 0x3FE5021930000000u, 0x3F2810C9u}, + {"6.56506150960922241210937501e-1", 0x3FE5021930000000u, 0x3F2810CAu}, + {"6565061509609222412109375000000000000000000001e-46", 0x3FE5021930000000u, 0x3F2810CAu}, + {"-0.6565061509609222412109375000000000000000000000000000000", 0xBFE5021930000000u, 0xBF2810CAu}, + {"6.5650615096092224e-1", 0x3FE5021930000000u, 0x3F2810C9u}, + {"-65650615096092225e-17", 0xBFE5021930000000u, 0xBF2810CAu}, + {"6565061509609222412e-19", 0x3FE5021930000000u, 0x3F2810C9u}, + {"0.6565061509609222413", 0x3FE5021930000000u, 0x3F2810CAu}, + {"0.65650615096092224121", 0x3FE5021930000000u, 0x3F2810C9u}, + {"6.5650615096092224122e-1", 0x3FE5021930000000u, 0x3F2810CAu}, + {"-6.5650615096092224121e-1", 0xBFE5021930000000u, 0xBF2810C9u}, + {"656506150960922241211e-21", 0x3FE5021930000000u, 0x3F2810CAu}, + {"18014627239033005156980393746124491372029297053813934326171875e-77", 0x3CA9F63970000000u, 0x254FB1CCu}, + {"0.00000000000000018014627239033005156980393746124491372029297053813934326171876", 0x3CA9F63970000000u, 0x254FB1CCu}, + {"-1.8014627239033005156980393746124491372029297053813934326171874e-16", 0xBCA9F63970000000u, 0xA54FB1CBu}, + {"1801462723903300515698039374612449137202929705381393432617187501e-79", 0x3CA9F63970000000u, 0x254FB1CCu}, + {"18014627239033005e-32", 0x3CA9F63970000000u, 0x254FB1CBu}, + {"0.00000000000000018014627239033006", 0x3CA9F63970000000u, 0x254FB1CCu}, + {"0.0000000000000001801462723903300515", 0x3CA9F63970000000u, 0x254FB1CBu}, + {"-1.801462723903300516e-16", 0xBCA9F63970000000u, 0xA54FB1CCu}, + {"-1.8014627239033005156e-16", 0xBCA9F63970000000u, 0xA54FB1CBu}, + {"18014627239033005157e-35", 0x3CA9F63970000000u, 0x254FB1CCu}, + {"180146272390330051569e-36", 0x3CA9F63970000000u, 0x254FB1CBu}, + {"-0.00000000000000018014627239033005157", 0xBCA9F63970000000u, 0xA54FB1CCu}, + {"0.000000000000000180146272390330051569803937461", 0x3CA9F63970000000u, 0x254FB1CBu}, + {"1.80146272390330051569803937462e-16", 0x3CA9F63970000000u, 0x254FB1CCu}, + {"0.05534819327294826507568359375", 0x3FAC569930000000u, 0x3D62B4CAu}, + {"-5.534819327294826507568359376e-2", 0xBFAC569930000000u, 0xBD62B4CAu}, + {"5534819327294826507568359374e-29", 0x3FAC569930000000u, 0x3D62B4C9u}, + {"-0.0553481932729482650756835937501", 0xBFAC569930000000u, 0xBD62B4CAu}, + {"5.534819327294826507568359375000000000000000000001e-2", 0x3FAC569930000000u, 0x3D62B4CAu}, + {"5534819327294826507568359375000000000000000000000000000000e-59", 0x3FAC569930000000u, 0x3D62B4CAu}, + {"0.055348193272948265", 0x3FAC569930000000u, 0x3D62B4C9u}, + {"5.5348193272948266e-2", 0x3FAC569930000000u, 0x3D62B4CAu}, + {"5.534819327294826507e-2", 0x3FAC569930000000u, 0x3D62B4C9u}, + {"5534819327294826508e-20", 0x3FAC569930000000u, 0x3D62B4CAu}, + {"55348193272948265075e-21", 0x3FAC569930000000u, 0x3D62B4C9u}, + {"-0.055348193272948265076", 0xBFAC569930000000u, 0xBD62B4CAu}, + {"0.0553481932729482650756", 0x3FAC569930000000u, 0x3D62B4C9u}, + {"-5.53481932729482650757e-2", 0xBFAC569930000000u, 0xBD62B4CAu}, + {"5.179692133247783258005389047985340416e36", 0x478F2C9450000000u, 0x7C7964A2u}, + {"5179692133247783258005389047985340417e0", 0x478F2C9450000000u, 0x7C7964A3u}, + {"5179692133247783258005389047985340415", 0x478F2C9450000000u, 0x7C7964A2u}, + {"-5.17969213324778325800538904798534041601e36", 0xC78F2C9450000000u, 0xFC7964A3u}, + {"-5179692133247783258005389047985340416000000000000000000001e-21", 0xC78F2C9450000000u, 0xFC7964A3u}, + {"-5179692133247783258005389047985340416.000000000000000000000000000000", 0xC78F2C9450000000u, 0xFC7964A2u}, + {"5.1796921332477832e36", 0x478F2C9450000000u, 0x7C7964A2u}, + {"51796921332477833e20", 0x478F2C9450000000u, 0x7C7964A3u}, + {"-5179692133247783258e18", 0xC78F2C9450000000u, 0xFC7964A2u}, + {"5179692133247783259000000000000000000", 0x478F2C9450000000u, 0x7C7964A3u}, + {"5179692133247783258000000000000000000", 0x478F2C9450000000u, 0x7C7964A2u}, + {"5.1796921332477832581e36", 0x478F2C9450000000u, 0x7C7964A3u}, + {"-5.179692133247783258e36", 0xC78F2C9450000000u, 0xFC7964A2u}, + {"-517969213324778325801e16", 0xC78F2C9450000000u, 0xFC7964A3u}, + {"517969213324778325800538904798e7", 0x478F2C9450000000u, 0x7C7964A2u}, + {"5179692133247783258005389047990000000", 0x478F2C9450000000u, 0x7C7964A3u}, + { + "0.22250738585072011360574097967091319759348195463516456480234261097248222220210769455165295239081350" + "8791414915891303962110687008643869459464552765720740782062174337998814106326732925355228688137214901" + "2981122451451889849057222307285255133155755015914397476397983411801999323962548289017107081850690630" + "6666559949382757725720157630626906633326475653000092458883164330377797918696120494973903778297049050" + "5108060994073026293712895895000358379996720725430436028407889577179615094551674824347103070260914462" + "1572289880258182545180325707018860872113128079512233426288368622321503775666622503982534335974568884" + "4239002654981983854879482922068947216898310996983658468140228542433306603398508864458040010349339704" + "2756718644338377048603786162277173854562306587467901408672332763671875e-307", 0x0010000000000000u, 0x00000000u + }, + { + "2.22507385850720113605740979670913197593481954635164564802342610972482222202107694551652952390813508" + "7914149158913039621106870086438694594645527657207407820621743379988141063267329253552286881372149012" + "9811224514518898490572223072852551331557550159143974763979834118019993239625482890171070818506906306" + "6665599493827577257201576306269066333264756530000924588831643303777979186961204949739037782970490505" + "1080609940730262937128958950003583799967207254304360284078895771796150945516748243471030702609144621" + "5722898802581825451803257070188608721131280795122334262883686223215037756666225039825343359745688844" + "2390026549819838548794829220689472168983109969836584681402285424333066033985088644580400103493397042" + "756718644338377048603786162277173854562306587467901408672332763671875000000000000000000001e-308", 0x0010000000000000u, 0x00000000u + }, + { + "0.11754942807573642917278829910357665133228589927589904276829631184250030649651730385585324256680905" + "8189392089843750000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "00000000000000000000000000000000000000000000000000000000000000000e-37", 0x380FFFFFE0000000u, 0x00800000u + }, + { + "1175494280757364291727882991035766513322858992758990427682963118425003064965173038558532425668090581" + "8939208984375000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000" + "00000000000001e-751", 0x380FFFFFE0000000u, 0x00800000u + }, + {"0", 0x0000000000000000u, 0x00000000u}, + {"-0", 0x8000000000000000u, 0x80000000u}, + {"0.0", 0x0000000000000000u, 0x00000000u}, + {"-0.0", 0x8000000000000000u, 0x80000000u}, + {"0e999999999999999999999", 0x0000000000000000u, 0x00000000u}, + {"-0.000e-99999", 0x8000000000000000u, 0x80000000u}, + {"1e-400", 0x0000000000000000u, 0x00000000u}, + {"-1e-400", 0x8000000000000000u, 0x80000000u}, + {"1e400", 0x7FF0000000000000u, 0x7F800000u}, + {"-1e400", 0xFFF0000000000000u, 0xFF800000u}, + {"1e-50", 0x358DEE7A4AD4B81Fu, 0x00000000u}, + {"-1e-50", 0xB58DEE7A4AD4B81Fu, 0x80000000u}, + {"1e39", 0x48078287F49C4A1Du, 0x7F800000u}, + {"-1e39", 0xC8078287F49C4A1Du, 0xFF800000u}, + {"1e99999999999999999999999999", 0x7FF0000000000000u, 0x7F800000u}, + {"1e-99999999999999999999999999", 0x0000000000000000u, 0x00000000u}, + {"1e0000000000000000000000000000000000000000308", 0x7FE1CCF385EBC8A0u, 0x7F800000u}, + {"123456789012345678901234567890e-30", 0x3FBF9ADD3746F65Fu, 0x3DFCD6EAu}, + {"18446744073709551615", 0x43F0000000000000u, 0x5F800000u}, + {"18446744073709551616", 0x43F0000000000000u, 0x5F800000u}, + {"-9223372036854775808", 0xC3E0000000000000u, 0xDF000000u}, + {"-9223372036854775809", 0xC3E0000000000000u, 0xDF000000u}, + } + }; + return table; +} + +} // namespace float_hard_cases diff --git a/tests/src/unit-class_lexer.cpp b/tests/src/unit-class_lexer.cpp index b56e4bd0b..c5e303c5d 100644 --- a/tests/src/unit-class_lexer.cpp +++ b/tests/src/unit-class_lexer.cpp @@ -13,15 +13,18 @@ using nlohmann::json; #include // array -#include // FLT_EVAL_METHOD #include // uint32_t, uint64_t +#include // snprintf #include // strtod #include // memcpy +#include // map #include // stringstream #include // string #include // pair #include // vector +#include "float_hard_cases.hpp" + namespace { // shortcut to scan a string literal @@ -257,7 +260,7 @@ TEST_CASE("lexer number fast path") "123456789012345678901234567890", // huge -> float "0.30000000000000004", "2.2250738585072014e-308", "1e308", // high-precision / wide-exponent values that exercise the - // std::from_chars (Eisel-Lemire) path beyond the Clinger subset + // Eisel-Lemire path beyond the Clinger subset "1.7976931348623157e308", "1.2345678901234567e-250", "9007199254740993", "5e-324", "1e-320" }; @@ -279,20 +282,18 @@ TEST_CASE("lexer number fast path") } } - SECTION("significant-digit gate for the Clinger fast path") + SECTION("significant digits around Clinger's fast path") { - // Clinger's fast path needs a significand below 2^53, so it cannot - // succeed once the mantissa has 17 or more significant digits (the - // significand would be at least 10^16). The lexer skips the attempt - // there. That is only allowed to save work: every value must still come - // out bit-exactly, and both scanners must agree. In particular the gate - // must not fire for tokens whose leading zeros merely look like extra - // digits - "0.1234567890123456" has 16 significant digits, not 17. + // Clinger's fast path needs a significand of at most 2^53, which + // tokens with 17 or more significant digits exceed. The conversion + // splits the token at the positions the scanners recorded, so leading + // zeros must not count as digits - "0.1234567890123456" has 16 + // significant digits, not 17 - and both scanners must agree. const std::vector numbers = { "1234567890123456", // 16 significant digits - "12345678901234567", // 17 -> attempt skipped - "123456789012345678", // 18 -> attempt skipped + "12345678901234567", // 17 + "123456789012345678", // 18 "0.1234567890123456", // 16: the leading "0" is not significant "0.12345678901234567", // 17 "0.00000000000000001", // 1, in a long token @@ -663,46 +664,145 @@ TEST_CASE("lexer string fast path") } } -TEST_CASE("parse_float_fast declines what it cannot convert exactly") +namespace { - // The lexer only hands well-formed numbers to parse_float_fast, so the - // malformed ones below can only be passed to it directly. Declining is - // always safe: the caller then falls back to a slower, exact conversion. - const auto fast = [](const std::string & s, double & out) +// the index of the decimal point (or npos) and of the end of the mantissa of a +// number token, which the lexer records while scanning it +std::pair float_token_layout(const std::string& s) +{ + std::size_t dot = std::string::npos; + std::size_t mantissa_end = s.size(); + for (std::size_t i = 0; i < s.size(); ++i) { - return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), out); - }; - double out = 0; + if (s[i] == '.') + { + dot = i; + } + else if (s[i] == 'e' || s[i] == 'E') + { + mantissa_end = i; + break; + } + } + return {dot, mantissa_end}; +} -#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0 - // without true double precision, the fast path declines everything - CHECK_FALSE(fast("1.5", out)); -#else - CHECK(fast("1.5", out)); - CHECK(out == 1.5); - CHECK(fast("+2.5e1", out)); - CHECK(out == 25.0); - CHECK(fast("-25E-1", out)); - CHECK(out == -2.5); - CHECK(fast("1e", out)); - CHECK(out == 1.0); -#endif +template +FloatType parse_native(const std::string& s) +{ + const auto layout = float_token_layout(s); + return nlohmann::detail::parse_float_native(s.data(), s.data() + s.size(), layout.first, layout.second); +} - // not a number - CHECK_FALSE(fast("", out)); - CHECK_FALSE(fast("-", out)); - CHECK_FALSE(fast(".", out)); - CHECK_FALSE(fast("1.2.3", out)); - CHECK_FALSE(fast("1x", out)); - CHECK_FALSE(fast("1e+", out)); - CHECK_FALSE(fast("1e1x", out)); +std::uint64_t bits_of(double d) +{ + std::uint64_t b = 0; + std::memcpy(&b, &d, sizeof(b)); + return b; +} - // numbers that are not represented exactly on the fast path - CHECK_FALSE(fast("12345678901234567890", out)); - CHECK_FALSE(fast("1e10000", out)); - CHECK_FALSE(fast("9007199254740993", out)); - CHECK_FALSE(fast("1e23", out)); - CHECK_FALSE(fast("1e-23", out)); +std::uint32_t bits_of(float f) +{ + std::uint32_t b = 0; + std::memcpy(&b, &f, sizeof(b)); + return b; +} + +std::uint64_t native_bits64(const std::string& s) +{ + return bits_of(parse_native(s)); +} + +std::uint32_t native_bits32(const std::string& s) +{ + return bits_of(parse_native(s)); +} +} // namespace + +TEST_CASE("parse_float_native rounds correctly") +{ + SECTION("double") + { + CHECK(native_bits64("1.5") == 0x3FF8000000000000u); + CHECK(native_bits64("0.1") == 0x3FB999999999999Au); + CHECK(native_bits64("-0.0") == 0x8000000000000000u); + CHECK(native_bits64("0e999999999999999999999") == 0u); + // 2^53 + 1 is exactly between two doubles: ties to even, unless more digits follow + CHECK(native_bits64("9007199254740993") == 0x4340000000000000u); + CHECK(native_bits64("9007199254740993.0000000000000000001") == 0x4340000000000001u); + CHECK(native_bits64("9007199254740992.9999999999999999999") == 0x4340000000000000u); + // 1 + 2^-53 exactly (a tie), and one unit in the 55th digit around it + CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203125") == 0x3FF0000000000000u); + CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203126") == 0x3FF0000000000001u); + CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203124") == 0x3FF0000000000000u); + // subnormal and overflow boundaries + CHECK(native_bits64("2.4703282292062327e-324") == 0u); + CHECK(native_bits64("2.4703282292062328e-324") == 1u); + CHECK(native_bits64("2.2250738585072011e-308") == 0x000FFFFFFFFFFFFFu); + CHECK(native_bits64("2.2250738585072012e-308") == 0x0010000000000000u); + CHECK(native_bits64("1.7976931348623157e308") == 0x7FEFFFFFFFFFFFFFu); + CHECK(native_bits64("1.7976931348623159e308") == 0x7FF0000000000000u); + CHECK(native_bits64("-1e400") == 0xFFF0000000000000u); + CHECK(native_bits64("-1e-400") == 0x8000000000000000u); + // exponents and zeros far beyond the range cancel out + CHECK(native_bits64("0." + std::string(1000, '0') + "1e1001") == 0x3FF0000000000000u); + CHECK(native_bits64("1" + std::string(1000, '0') + "e-1000") == 0x3FF0000000000000u); + CHECK(native_bits64("1e-99999999999999999999999") == 0u); + CHECK(native_bits64("1E+99999999999999999999999") == 0x7FF0000000000000u); + // more digits than any midpoint has (769): only whether a nonzero digit follows matters + const std::string tie = "1.00000000000000011102230246251565404236316680908203125"; + CHECK(native_bits64(tie + std::string(800, '0')) == 0x3FF0000000000000u); + CHECK(native_bits64(tie + std::string(800, '0') + "1") == 0x3FF0000000000001u); + } + + SECTION("float") + { + CHECK(native_bits32("1.5") == 0x3FC00000u); + CHECK(native_bits32("0.1") == 0x3DCCCCCDu); + CHECK(native_bits32("-0.0") == 0x80000000u); + // 2^24 + 1 is exactly between two floats + CHECK(native_bits32("16777217") == 0x4B800000u); + CHECK(native_bits32("16777217.000000000000000000001") == 0x4B800001u); + CHECK(native_bits32("16777218.999999999999999999999") == 0x4B800001u); + CHECK(native_bits32("16777219") == 0x4B800002u); + // subnormal and overflow boundaries + CHECK(native_bits32("3.4028235677973366e38") == 0x7F7FFFFFu); + CHECK(native_bits32("3.4028235677973367e38") == 0x7F800000u); + CHECK(native_bits32("7.006492321624085e-46") == 0u); + CHECK(native_bits32("7.006492321624086e-46") == 1u); + CHECK(native_bits32("1.1754942e-38") == 0x007FFFFFu); + CHECK(native_bits32("-1.17549435e-38") == 0x80800000u); + CHECK(native_bits32("1e39") == 0x7F800000u); + CHECK(native_bits32("-1e-50") == 0x80000000u); + // not rounded through double: its double would round to another float + CHECK(native_bits32("1.00000005960464477539062500000000001") == 0x3F800001u); + CHECK(native_bits32("9007199254740993") == 0x5A000000u); + } + + SECTION("the conversion shared with other parsers") + { + // convert_float() gives the lexer's results, for every type + const std::vector tokens = + { + "0", "-0.0", "1.5", "0.1", "1e-400", "-2.5E+3", "123456789012345678901234567890", + "9007199254740993.0000000000000000001", "4.9406564584124654e-324" + }; + using float_json = nlohmann::basic_json; + using long_double_json = nlohmann::basic_json; + for (const auto& t : tokens) + { + CAPTURE(t); + const auto layout = float_token_layout(t); + const char* const first = t.data(); + const char* const last = first + t.size(); + const auto d = nlohmann::detail::convert_float(first, last, layout.first, layout.second); + const auto f = nlohmann::detail::convert_float(first, last, layout.first, layout.second); + const auto ld = nlohmann::detail::convert_float(first, last, layout.first, layout.second); + CHECK(bits_of(d) == bits_of(json::parse(t).get())); + CHECK(bits_of(f) == bits_of(float_json::parse(t).get())); + CHECK(ld == long_double_json::parse(t).get()); + } + } } namespace @@ -806,40 +906,6 @@ std::size_t big_bit_length(const big_uint& a) } return n; } - -std::uint64_t bits_of(double d) -{ - std::uint64_t b = 0; - std::memcpy(&b, &d, sizeof(b)); - return b; -} - -bool eisel_lemire(const std::string& s, double& out) -{ - return nlohmann::detail::parse_float_eisel_lemire(s.data(), s.data() + s.size(), out); -} - -// significant digits of a token, without trailing zeros -std::size_t significant_digits(const std::string& s) -{ - std::string digits; - for (const char c : s) - { - if (c == 'e' || c == 'E') - { - break; - } - if (c >= '0' && c <= '9' && !(digits.empty() && c == '0')) - { - digits += c; - } - } - while (!digits.empty() && digits.back() == '0') - { - digits.pop_back(); - } - return digits.size(); -} } // namespace TEST_CASE("Eisel-Lemire float conversion") @@ -1237,26 +1303,33 @@ TEST_CASE("Eisel-Lemire float conversion") for (const auto& c : known) { CAPTURE(c.first); - double out = 0; - if (eisel_lemire(c.first, out)) - { - CHECK(bits_of(out) == c.second); - } - else - { - // only tokens with more than 19 significant digits are left to - // strtod: those whose value lies too close to a tie - CHECK(significant_digits(c.first) > 19); - } + CHECK(native_bits64(c.first) == c.second); } } + SECTION("binary32") + { + using binary32 = nlohmann::detail::ieee_binary_format<24>; + CHECK(nlohmann::detail::eisel_lemire(0, 1) == 0x3F800000u); + CHECK(nlohmann::detail::eisel_lemire(-1, 1) == 0x3DCCCCCDu); + CHECK(nlohmann::detail::eisel_lemire(-1, 15) == 0x3FC00000u); + CHECK(nlohmann::detail::eisel_lemire(0, 16777217) == 0x4B800000u); // tie, to even + CHECK(nlohmann::detail::eisel_lemire(0, 16777219) == 0x4B800002u); // tie, to even + CHECK(nlohmann::detail::eisel_lemire(-45, 1) == 0x00000001u); + CHECK(nlohmann::detail::eisel_lemire(-46, 7) == 0x00000000u); + CHECK(nlohmann::detail::eisel_lemire(-46, 8) == 0x00000001u); + CHECK(nlohmann::detail::eisel_lemire(-65, 9999999999999999999u) == 0x00000000u); + CHECK(nlohmann::detail::eisel_lemire(20, 3402823466385288598u) == 0x7F7FFFFFu); + CHECK(nlohmann::detail::eisel_lemire(20, 3402823669209384635u) == 0x7F800000u); + CHECK(nlohmann::detail::eisel_lemire(39, 1) == 0x7F800000u); + CHECK(nlohmann::detail::eisel_lemire(-5, 0) == 0x00000000u); + } + SECTION("round trip") { - // every double written by to_chars and read back, also with trailing - // digits that make the token longer than 19 digits + // every double written by to_chars and read back, and its 17-digit + // form with trailing digits that make the token longer than 19 digits std::uint64_t state = 5295; - std::size_t declined = 0; for (int i = 0; i < 200000; ++i) { state ^= state << 13u; @@ -1278,30 +1351,51 @@ TEST_CASE("Eisel-Lemire float conversion") const char* end = nlohmann::detail::to_chars(buffer.data(), buffer.data() + buffer.size(), d); const std::string token(buffer.data(), static_cast(end - buffer.data())); CAPTURE(token); - double out = 0; - REQUIRE(eisel_lemire(token, out)); - CHECK(bits_of(out) == b); + CHECK(native_bits64(token) == b); - // insert digits before the exponent: the value moves by far less - // than the distance to the rounding boundary, so it must not change - std::string longer = token; + // insert digits before the exponent of the 17-digit form: that + // form lies strictly inside the rounding interval of the double + // (the shortest one may lie on its boundary), and the digits move + // it by far less than the distance to the boundary, so the value + // must not change + std::array digits17{}; + static_cast(std::snprintf(digits17.data(), digits17.size(), "%.17g", d)); // NOLINT(cppcoreguidelines-pro-type-vararg,hicpp-vararg) + std::string longer = digits17.data(); const std::size_t e = longer.find('e'); const std::size_t dot = longer.find('.'); const std::string extra = dot == std::string::npos ? ".000000000000000000001" : "000000000000000000001"; longer.insert(e == std::string::npos ? longer.size() : e, extra); CAPTURE(longer); - if (eisel_lemire(longer, out)) - { - CHECK(bits_of(out) == b); - } - else - { - // w and w + 1 round differently: only when the value is very - // close to a rounding boundary - ++declined; - } + CHECK(native_bits64(longer) == b); + } + } + + SECTION("round trip, binary32") + { + std::uint32_t state = 5295; + for (int i = 0; i < 100000; ++i) + { + state ^= state << 13u; + state ^= state >> 17u; + state ^= state << 5u; + std::uint32_t b = state; + if ((b & 0x7F800000u) == 0x7F800000u) + { + continue; // infinity or NaN + } + if (i % 4 == 0) + { + b &= 0x807FFFFFu; // subnormals + } + float f = 0; + std::memcpy(&f, &b, sizeof(f)); + + std::array buffer{}; + const char* end = nlohmann::detail::to_chars(buffer.data(), buffer.data() + buffer.size(), f); + const std::string token(buffer.data(), static_cast(end - buffer.data())); + CAPTURE(token); + CHECK(native_bits32(token) == b); } - CHECK(declined < 1000); // 107 of the 200,000 } SECTION("used by the lexer") @@ -1315,3 +1409,76 @@ TEST_CASE("Eisel-Lemire float conversion") "[json.exception.out_of_range.406] number overflow parsing '1.7976931348623159e308'", json::out_of_range&); } } + +namespace +{ +using float_json = nlohmann::basic_json; + +// the bits of the float that parse() gives for a token, via both scanners; +// the value must be the same for both +template +void check_parse(const std::string& token, Bits expected, Bits infinity) +{ + std::stringstream stream(token); + if ((expected & ~(Bits{1} << (8 * sizeof(Bits) - 1))) == infinity) + { + Json _; + CHECK_THROWS_WITH_AS(_ = Json::parse(token), ("[json.exception.out_of_range.406] number overflow parsing '" + token + "'").c_str(), typename Json::out_of_range&); + CHECK_THROWS_WITH_AS(_ = Json::parse(stream), ("[json.exception.out_of_range.406] number overflow parsing '" + token + "'").c_str(), typename Json::out_of_range&); + return; + } + const Json contiguous = Json::parse(token); + const Json streamed = Json::parse(stream); + if (contiguous.is_number_float()) // not an integer that fits + { + CHECK(bits_of(contiguous.template get()) == expected); + CHECK(bits_of(streamed.template get()) == expected); + } + else + { + CHECK(streamed.is_number_integer()); + } +} +} // namespace + +TEST_CASE("float conversion of hard cases") +{ + // see float_hard_cases.hpp + for (const auto& c : float_hard_cases::cases()) + { + const std::string token = c.token; + CAPTURE(token); + CHECK(native_bits64(token) == c.bits64); + CHECK(native_bits32(token) == c.bits32); + check_parse(token, c.bits64, std::uint64_t{0x7FF0000000000000u}); + check_parse(token, c.bits32, std::uint32_t{0x7F800000u}); + } +} + +TEST_CASE("float overflow and underflow in the parser") +{ + SECTION("double") + { + check_parse("1.7976931348623157e308", std::uint64_t{0x7FEFFFFFFFFFFFFFu}, std::uint64_t{0x7FF0000000000000u}); + check_parse("1.7976931348623159e308", std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u}); + check_parse("-1e309", std::uint64_t{0xFFF0000000000000u}, std::uint64_t{0x7FF0000000000000u}); + check_parse("1" + std::string(400, '0'), std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u}); + check_parse("1e99999999999999999999", std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u}); + // an underflow gives a zero with the sign of the token + check_parse("1e-400", std::uint64_t{0}, std::uint64_t{0x7FF0000000000000u}); + check_parse("-1e-400", std::uint64_t{0x8000000000000000u}, std::uint64_t{0x7FF0000000000000u}); + check_parse("-2.4703282292062327e-324", std::uint64_t{0x8000000000000000u}, std::uint64_t{0x7FF0000000000000u}); + check_parse("0." + std::string(400, '0') + "1", std::uint64_t{0}, std::uint64_t{0x7FF0000000000000u}); + } + + SECTION("float") + { + check_parse("3.4028234e38", std::uint32_t{0x7F7FFFFFu}, std::uint32_t{0x7F800000u}); + check_parse("3.4028236e38", std::uint32_t{0x7F800000u}, std::uint32_t{0x7F800000u}); + check_parse("-1e39", std::uint32_t{0xFF800000u}, std::uint32_t{0x7F800000u}); + check_parse("1e-46", std::uint32_t{0}, std::uint32_t{0x7F800000u}); + check_parse("-1e-46", std::uint32_t{0x80000000u}, std::uint32_t{0x7F800000u}); + check_parse("-7.006492321624085e-46", std::uint32_t{0x80000000u}, std::uint32_t{0x7F800000u}); + check_parse("-7.006492321624086e-46", std::uint32_t{0x80000001u}, std::uint32_t{0x7F800000u}); + } +} diff --git a/tests/src/unit-locale-cpp.cpp b/tests/src/unit-locale-cpp.cpp index 14f743a66..eeddda358 100644 --- a/tests/src/unit-locale-cpp.cpp +++ b/tests/src/unit-locale-cpp.cpp @@ -257,10 +257,11 @@ struct LocaleSwitchingSax final: public nlohmann::json_sax TEST_CASE("locale changes between lexer construction and number conversion (#5198)") { - // The numbers are chosen so that the conversion also takes the strtod - // fallback, which honors the locale that is current at conversion time: - // too many significant digits for Clinger's fast path, an underflow that - // std::from_chars rejects, and a plain value. + // float and double are converted without the locale. A long double that + // is not binary64 can take the strtold fallback, which honors the locale + // that is current at conversion time. The numbers are chosen so that it + // does: too many significant digits for Clinger's fast path, an underflow + // that std::from_chars rejects, and a plain value. const std::vector numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"}; std::string text = "["; for (const auto& n : numbers) @@ -324,7 +325,8 @@ TEST_CASE("locale changes between lexer construction and number conversion (#519 } } - // a long double goes through std::strtold unless std::from_chars supports it + // a long double goes through std::strtold unless it is binary64 or + // std::from_chars supports it { bool switched = false; const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept @@ -350,8 +352,9 @@ TEST_CASE("locale with a multi-byte decimal point") { // Some locales use a decimal point that is not a single character, e.g. // U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be - // substituted in place for '.', so the strtod fallback stops early. The - // conversion must still terminate rather than retry forever. + // substituted in place for '.', so the strtold fallback (only for long + // double formats other than binary64) stops early. The conversion must + // still terminate rather than retry forever. const std::array names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}}; bool tested = false; for (const char* name : names) @@ -369,12 +372,21 @@ TEST_CASE("locale with a multi-byte decimal point") tested = true; // too many significant digits for Clinger's fast path, and an underflow - // that std::from_chars rejects: both reach the strtod fallback + // that std::from_chars rejects: double does not depend on the locale json j; CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]")); CHECK(j.is_array()); + CHECK(j[0] == 3.14159265358979323846); + CHECK(j[1] == 0.0); + CHECK(j[2] == -0.000123456789012345678); CHECK(json::accept("3.14159265358979323846")); + // a long double that reaches the strtold fallback must still terminate + using long_double_json = nlohmann::basic_json; + long_double_json ld; + CHECK_NOTHROW(ld = long_double_json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]")); + CHECK(ld.is_array()); + // a value the locale-independent paths convert is not affected CHECK(json::parse("12.5") == 12.5); }