mirror of
https://github.com/nlohmann/json.git
synced 2026-10-01 22:45:17 +00:00
* Drop stale LCOV_EXCL_LINE from the json_pointer out_of_range.410 throw The comment said the size_type overflow check in array_index() is only triggered on special platforms like 32-bit, and the throw was excluded from coverage. On 64-bit platforms the check is true for SIZE_MAX itself, and unit-json_pointer.cpp has asserted that case four times since #5395, so the line is executed in the coverage job. Reword the comment and remove the exclusion marker so the coverage report notices if the tests stop reaching it. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name all three C-array check aliases in the enum-macro NOLINTs NLOHMANN_JSON_SERIALIZE_ENUM(_STRICT) suppressed the c-array warning under modernize-avoid-c-arrays only, but clang-tidy emits the same diagnostic under the aliases cppcoreguidelines-avoid-c-arrays and hicpp-avoid-c-arrays too. Any user running those checks got a false positive at every macro expansion, and our own tests needed a local NOLINT at each call site to work around it. Name all three aliases in the four macro comments instead, and drop the now-redundant c-array names from the five test call-site NOLINTs. Comment-only change; behavior, the public API, and the ABI do not change. Ran make amalgamate. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Include doctest as a SYSTEM directory instead of disabling warnings for all tests test_main added -Wno-deprecated and -Wno-float-equal as PUBLIC compile options for every non-MSVC compiler, so they were applied to every translation unit, library headers included, and silenced the CI warnings meant to check the library's own -Wfloat-equal pragmas. The only code that actually needed the suppression was the vendored doctest.h, which was included as a normal (non-SYSTEM) directory. Include thirdparty/doctest as SYSTEM for test_main, matching what tests/abi/CMakeLists.txt already does, and drop the two suppressions from both targets. Verified locally that unit-comparison, unit-conversions and unit-constructor1 compile clean with -Werror -Weverything and doctest as -isystem, and that CMake still configures with JSON_BuildTests=ON. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the no-op ci_clang_analyze target ci_clang_analyze configured the build with the real compiler and only then wrapped ninja with scan-build. scan-build intercepts compiles by overriding CC/CXX, but build.ninja already had the compiler path baked in from the configure step, so every run bypassed the analyzer: CI logs show "No bugs found" after a normal build, never an analysis. The job also used Debian's frozen clang-tools-14 rather than the image's own clang, and CLANG_ANALYZER_CHECKS still named three valist.* checkers that current clang merged into security.VAList. ci_clang_tidy already runs every clang-analyzer-* check (via .clang-tidy's "Checks: '*'") with warnings as errors, so nothing is lost by removing the dead job. Delete ci_clang_analyze, CLANG_ANALYZER_CHECKS and the SCAN_BUILD_TOOL lookup from cmake/ci.cmake, drop it from the ubuntu.yml ci_static_analysis_clang matrix, and drop the now-unused clang-tools apt package (iwyu stays for ci_single_binaries). Reword quality_assurance.md and assurance_case.md, which described the dead job as a working control, to say the Clang Static Analyzer checks run through clang-tidy. Verified that `cmake -DJSON_CI=ON` still configures cleanly and that ci_clang_analyze no longer appears in the generated build or in any CMake/workflow file. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Re-enable portability-template-virtual-member-function; remove redundant forwards .clang-tidy disabled three checks "to get the CI going" (#4489, 2024-11-13): portability-template-virtual-member-function, bugprone-use-after-move and its alias hicpp-invalid-access-moved. portability-template-virtual-member-function only flagged output_stream_adapter::write_character/write_characters; annotate both with NOLINT and re-enable the check. bugprone-use-after-move flagged several double forwards that have no effect at runtime: - from_json.hpp calls std::forward<BasicJsonType>(j).at(Idx) inside pack expansions; at() has no ref-qualified overloads and always returns an lvalue reference, so the forward is a no-op. Replace with plain j.at(Idx) in all four places. - the move constructor forwards the whole object to its base class and then reads other's members. That is item 9 of #5724 (together with its cppcheck suppressions) and is left to that change. - input_adapters.hpp forwards the container twice on purpose, so the begin/end iterator types match adapter_type; annotate with NOLINT and a comment instead of changing behavior. The check still flags the move constructor (see above) and two sites in at(KeyType&&) (both overloads, json.hpp, in the throw's string_t(std::forward<KeyType>(key)) after find(std::forward<KeyType>(key))). Open PR #5689 rewrites that hunk, so bugprone-use-after-move (and hicpp-invalid-access-moved) stay disabled for now, with a comment explaining why; re-enable them once #5689 and the #5724 move-constructor change have landed. Also resolve the portability-avoid-pragma-once TODO: single_include never has #pragma once (amalgamate.py strips it) and every supported compiler accepts it in include/, so keep it disabled with an explanatory comment instead of a TODO. Fix the stale "json.hpp, around line 1265" comment in unit-class_parser.cpp, which now points at the move constructor's actual line. Behavior, the public API and the ABI do not change. Verified with clang-tidy 22.1.8 that portability-template-virtual-member-function now reports nothing, that bugprone-use-after-move/ hicpp-invalid-access-moved report only the known at(KeyType&&) and move-constructor sites, and that unit-custom-base-class, unit-constructor1, unit-conversions, unit-element_access2, unit-class_parser and unit-diagnostic-positions (JSON_DIAGNOSTIC_POSITIONS=1) compile under ASan/UBSan and pass with the same assertion counts as before. Ran make amalgamate. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale and malformed NOLINT comments json_sax.hpp named "-warnings-as-errors" in the NOLINT list on the two JSON_ASSERT(false) lines; that is the suffix clang-tidy appends to a diagnostic tag under WarningsAsErrors, not a check name, and every other JSON_ASSERT(false) omits it. unit-capacity.cpp carried 30 "// NOLINT(misc-const-correctness)" comments on "json j = ...;" declarations that are all used with non-const members afterwards, so the check has nothing to report there. unit-constructor2.cpp used a blanket "// NOLINT: access after move is OK here" on a use-after-move that hides every check on the line; naming bugprone-use-after-move and hicpp-invalid-access-moved keeps the intent once those checks are re-enabled (#5724). Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 10 * Remove stale .clang-tidy entries -google-runtime-references disabled a check that neither clang-tidy 22.1.8 nor 23.1.2 lists under --list-checks -checks='*'; it was removed upstream. The commented-out HeaderFilterRegex line has been unused since the active HeaderFilterRegex was introduced in #2561 (2021). Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 11 * Remove the GCC C++20 -Wignored-attributes pragma in json.hpp The pragma (added in #5164) claimed to work around the C++ modules redefinition errors of #5103, but #5103 is about hard errors (e.g. "redefinition of std::__is_constant_evaluated()", conflicting std::integral_constant) that ignoring a warning cannot suppress; they are traced to GCC PR 124430 and reproduce with <map> or <string> instead of json.hpp too. A GCC 16.2 -std=gnu++20 -fmodules build following #5103's repro steps still fails with the pragma in place, and a build of all test TUs with GCC_CXXFLAGS (which enable -Wignored-attributes) and the pragma removed produces no such warning. The block only hid a warning class from GCC C++20 users while suggesting #5103 was handled. Overlaps #5610, whose hunks touch the closing half of this pragma to insert the json_literals.hpp include. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 9 * Fix stale doxygen comments hidden by the -Wdocumentation pragma macro_scope.hpp ignores -Wdocumentation and -Wdocumentation-unknown-command for the whole library, which also hides genuine documentation mistakes: - detail::unescape() documented "@return unescaped string" but returns void and unescapes its argument in place; reworded to "@param[in,out] s string to unescape in place" and dropped the bogus @return. - basic_json::get()'s copy-conversion overload wrote "converted to @tparam ValueType" inside @return, which Doxygen and Clang parse as a second, malformed @tparam; changed to "@a ValueType", matching the two other get() overloads a few lines above that already use it. This narrows the gap the -Wdocumentation pragma needs to cover; fully replacing the Doxygen-only commands it also hides (item 2c) is left for after #5267. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 2 * Fix -Wextra-semi-stmt at its actual source, not assert() clang_flags.cmake blamed the global -Wno-extra-semi-stmt on assert(), but assert() expands to an expression under glibc and libc++ and does not trigger this warning. unit-assert_macro.cpp overrides JSON_ASSERT with "{if (!(x)) ++assert_counter; }", a bare block followed by a semicolon at every JSON_ASSERT(...) call site in the library; that was the actual source of 151 of the 208 -Wextra-semi-stmt sites found in a Clang 22 -Weverything sweep of the test suite with the flag removed. Switched to the standard do/while(false) macro idiom, which does not expand to a statement-plus-semicolon, and corrected the comment to name the remaining source instead: vendored Doctest's CAPTURE(x) shim, which already ends in a semicolon. Verified with clang++ -Wextra-semi-stmt (plus the file's other CI ignores) that unit-assert_macro.cpp now compiles without any -Wextra-semi-stmt diagnostic. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 8 (step 1 of 2; step 2 covers the CAPTURE() call sites) * Drop the redundant semicolon from CAPTURE() call sites; remove -Wno-extra-semi-stmt doctest_compatibility.h defines CAPTURE(x) as DOCTEST_CAPTURE(x); (with a trailing semicolon baked into the macro), specifically so call sites do not need to add one themselves; most of the ~267 call sites already follow that convention. The remaining 64 call sites across 20 files wrote "CAPTURE(x);" anyway, turning into a statement plus an empty statement and triggering -Wextra-semi-stmt. Dropped the redundant semicolon at each of those sites. With item 6 having already made vendored Doctest a SYSTEM include, and this the last known source of -Wextra-semi-stmt findings, removed the flag from clang_flags.cmake entirely. Verified with clang++ -Wextra-semi-stmt (plus the file's other CI ignores) that all 20 touched files, plus a file with no CAPTURE() use (unit-json_pointer.cpp), compile without any -Wextra-semi-stmt diagnostic. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 8 (step 2 of 2) * Switch ci_static_analysis_clang off the frozen LLVM 22 dev image ubuntu.yml pinned the clang-tidy/clang-tidy-sanitizer/single-binaries job to silkeh/clang:dev, a tag last pushed 2026-02-18 that reports "clang version 22.0.0 (...+20251015...)", a pre-release snapshot from before the LLVM 22 release; the maintainer now updates dev-unstable, 22, and latest instead. Switched to silkeh/clang:22, matching the other clang jobs on :latest. Verified with clang-tidy 22.1.8 (the image's actual version) against this repository's .clang-tidy and library headers what the release image newly reports compared to :dev: - readability-redundant-typename fires at ~250 sites across the _cpp20-relevant conversion/to_chars headers; the library targets C++11 and keeps the typenames, so the check is disabled in .clang-tidy, matching how the file already handles checks that don't fit a C++11 codebase. - misc-anonymous-namespace-in-header fires on the two anonymous namespaces in from_json.hpp and to_json.hpp; added the alias to their existing NOLINT (cert-dcl59-cpp, fuchsia-header-anon-namespaces, google-build-namespaces). - bugprone-std-namespace-modification fires on every addition to namespace std: the std::hash, std::formatter and std::swap overloads in json.hpp, and the std::tuple_size/std::tuple_element specializations in iteration_proxy.hpp (this last file is not named in #5725's item 5, found by actually running clang-tidy 22.1.8 against the current tree). All six are legal, deliberate additions to namespace std (explicit/partial specializations of std types, or the pre-C++20 std::swap overload); annotated each with the check name next to its existing cert-dcl58-cpp NOLINT. - modernize-avoid-c-style-cast reported nothing new. Also added clang++-22/21, clang-tidy-22/21, g++-16 and gcov-16 to the find_program search lists in ci.cmake so a local "maximal warnings" configure prefers the current toolchain version over an older one on PATH. #5725 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Regenerate cmake/gcc_flags.cmake for GCC 16.2.0 GCC_CXXFLAGS was generated for GCC 15.1.0, but ci_test_gcc and ci_test_gcc_cxx{11..26} now run in gcc:latest, currently GCC 16.2.0, so the "maximal warnings" job was missing warnings introduced since 15.1.0 while carrying entries GCC 16 treats as duplicates or no-ops. Regenerated with https://github.com/nlohmann/gcc_flags (patched locally to not crash on an option whose "-x c++ <opt> -" probe fails before it reads stdin, e.g. -Wabi=; the tool otherwise raises BrokenPipeError instead of recording the option as an error) run against g++ 16.2.0 in the official gcc:16 Docker image, keeping the documented -Wno-* exclusions and the same alphabetical placement scheme as before. Also added three GCC 16 warnings the generator cannot discover on its own because it only probes value ranges/lists it finds in the -Q option name itself, not in the enum choices --help=warnings documents separately: - -Wbidi-chars=any, -Wleading-whitespace=spaces: manually verified these compile cleanly with g++ 16.2.0. - -Wstrict-flex-arrays: deliberately NOT added, unlike the other two. Without -fstrict-flex-arrays (which the library does not enable, as it would change codegen for flexible array members), GCC prints "'-Wstrict-flex-arrays' is ignored when '-fstrict-flex-arrays' is not present" on every translation unit, and under our -Werror that note itself aborts the build. This differs from the harmless no-op warnings already kept in the file (-Whsa, -Wsynth, -Wunreachable-code, -Wunsafe-loop-optimizations), which emit nothing; #5725 item 7 named -Wstrict-flex-arrays as one of the flags GCC 16 adds, but did not anticipate this failure mode. Verified: compiled the library header and a representative set of test translation units (including ones touched by items 1, 3, 8, 9, 10 of this issue) with the regenerated GCC_CXXFLAGS plus -Werror under g++ 16.2.0 at -std=c++11 through -std=c++26, with zero warnings; ran the full local test suite (129/129 passing, unrelated to this compiler) as a regression check. CI must still confirm the actual ci_test_gcc / ci_test_standards_gcc targets end to end, since this was verified with direct g++ invocations rather than through the CMake/ CXXFLAGS environment-variable plumbing in ci.cmake. #5725 item 7 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Avoid std::basic_string<CharType> for non-character output_adapter CharType output_adapter<CharType, StringType> defaulted StringType to std::basic_string<CharType>, and (with JSON_NO_IO undefined) always declared a std::basic_ostream<CharType>&-taking constructor. For CharType with no non-deprecated std::char_traits specialization (only std::uint8_t is ever used this way, by the binary writers), simply naming either type - as an unused default template argument, or as an unused, never-called constructor's parameter type - instantiates std::char_traits<CharType> merely to name it, which some standard libraries mark deprecated: with the library-wide -Wdocumentation pragma (item 2's other half, left for a later commit) temporarily removed, an Apple clang 21 / libc++ TU calling json::to_cbor(j, vec) with std::vector<std::uint8_t>& got one -Wdeprecated-declarations warning per binary writer at the old output_adapters.hpp:193. Replaced the eager std::basic_string<CharType> / std::basic_ostream <CharType> defaults with a bool-tagged partial specialization (not std::conditional, which requires naming both branches' types up front regardless of which is selected, reproducing the same warning) that only ever names std::basic_string<CharType> / std::basic_ostream <CharType> when CharType is actually one of char, wchar_t, char16_t, char32_t, or (with __cpp_lib_char8_t) char8_t. For any other CharType, output_adapter's StringType and ostream-constructor parameter fall back to two distinct empty placeholder types, kept distinct so the two constructor overloads do not collide into a single redeclaration. Public API / behavior: passing a std::basic_string<std::uint8_t>& or std::basic_ostream<std::uint8_t>& directly to a binary writer's output_adapter now fails to compile instead of compiling with a deprecation warning; this was neither documented nor tested. All documented uses (std::vector<CharType>, std::basic_ostream<CharType> and StringType for character CharType) are unaffected. Verified with Apple clang 21 / libc++, with the two -Wdocumentation* "ignored" pragma lines in macro_scope.hpp temporarily removed and -std=c++11/c++20 plus the project's -Weverything flag set: calling to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson/to_bon8 on a std::vector<std::uint8_t> now produces no char_traits<unsigned char> (or any other) deprecation warning, while the char-based string- and ostream-adapter paths, and a to_cbor/from_cbor round trip, still compile and run correctly; also verified with GCC 16.2.0. Ran the full local test suite, including the binary-format unit tests (unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bson, unit-bon8, unit-binary_writer_sinks, unit-binary_formats, unit-custom-binary-type): 129/129 passing. #5725 item 2 (step a) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the library-wide -Wdocumentation pragma; fix what it hid macro_scope.hpp / macro_unscope.hpp pushed and popped a Clang diagnostic region over the entire library that ignored -Wdocumentation and -Wdocumentation-unknown-command. Removed both pragmas and fixed every finding a full -Wdocumentation (which implies -Wdocumentation-unknown-command and -Wdocumentation-deprecated-sync) build reports, so the library now compiles clean under Clang's documentation checks without a blanket suppression. Overlaps #5267, which is still open and edits a nearby doc block (json.hpp's get()/get_impl() @return, already fixed in the item 2 step (b) commit of this branch); this commit does not touch that block again. Unknown Doxygen alias commands (Doxyfile removed in #3071, so these were never rendered by anything) rewritten as plain prose, keeping the same information: - @requirement REQ-JSON-01 / REQ-JSON-02 (iter_impl.hpp, json_reverse_iterator.hpp): now "This class satisfies the following concept requirements (REQ-JSON-0N):". - @liveexample{prose,example-id} (three sites in json.hpp): kept the prose, dropped the command wrapper and the trailing example-id (docs/mkdocs/docs/examples/*.cpp still exist and are used directly by the rendered docs, not through this in-header alias) and unescaped the "\," commas that were only needed for the old alias's comma-separated argument syntax. - @complexity X (json.hpp x4, json_pointer.hpp x2, serializer.hpp x1): now "Complexity: X". Backslash sequences Clang's comment lexer tried to parse as commands, escaped to render as literal backslashes: - lexer.hpp get_codepoint(): two `\u` occurrences. - binary_reader.hpp get_bson_cstr() / get_bson_cstr_bulk(): two `\x00` occurrences. - serializer.hpp: three `\uXXXX` occurrences (constructor @param, append_codepoint_to_string_buffer() @brief, and the ensure_ascii member comment). One finding remained after all of the above: Clang reports "declaration is marked with '@deprecated' command but does not have a deprecation attribute" on the deprecated sax_parse(span_input_adapter&&, ...) overload, even though JSON_HEDLEY_DEPRECATED_FOR does expand to __attribute__((deprecated(...))) for Clang. Several isolated reproductions of this exact declaration shape - doc comment, template<>, two stacked __attribute__ macros, an overload set sharing the name - did not reproduce the warning, so this looks like a Clang comment/declaration-association quirk specific to this overload inside the much larger basic_json class template, not an actual documentation defect. Rather than keep the pragma library-wide for one Clang false positive, added a tightly scoped -Wdocumentation-deprecated-sync push/pop around just that overload. Verified with Apple clang 21 and the project's actual -Weverything flag set (cmake/clang_flags.cmake) on the full header at -std=c++11 and -std=c++20: zero -Wdocumentation* diagnostics. Also compiled clean with GCC 16.2.0 (the pragmas are already __clang__-gated, so this only confirms no unrelated breakage). Ran make check-amalgamation and the full local test suite: 129/129 passing. #5725 item 2 (step c) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Take the JSON value by const reference in the array and tuple from_json paths Review feedback on #5737 (gregmarr): once the no-op std::forward calls are gone, the forwarding references have no purpose. from_json_fn passes the value as const BasicJsonType&, so these functions were only ever instantiated with a const lvalue anyway. The std::array, std::pair and std::tuple overloads of from_json and their helpers now take const BasicJsonType& and pass j on unchanged. Because the deduced BasicJsonType is now the plain type, tuple_type and the static_assert name const BasicJsonType& explicitly, so the reference checks are unchanged: get<std::tuple<const std::string&>>() still works, and get<std::tuple<std::string&>>() still fails the same static_assert. from_json_tuple_get_impl keeps its forwarding reference, since tuple_type calls it through std::declval. Behavior, the public API and the ABI do not change. unit-conversions, unit-constructor1, unit-udt, unit-udt_macro, unit-regression1/2/3, unit-deserialization, unit-noexcept, unit-items, unit-allocator, unit-custom-object-type, unit-ordered_json2 and unit-brace-init-copy-semantics pass at C++11, C++17 and C++20 with unchanged assertion counts. Ran make amalgamate. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
898 lines
36 KiB
C++
898 lines
36 KiB
C++
// __ _____ _____ _____
|
|
// __| | __| | | | JSON for Modern C++
|
|
// | | |__ | | | | | | version 3.12.0
|
|
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
|
|
//
|
|
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
|
|
// SPDX-License-Identifier: MIT
|
|
|
|
#pragma once
|
|
|
|
#include <algorithm> // min
|
|
#include <array> // array
|
|
#include <cstddef> // size_t
|
|
#include <cstdint> // uint32_t
|
|
#include <cstring> // strlen
|
|
#include <iterator> // begin, end, iterator_traits, random_access_iterator_tag, distance, next
|
|
#include <streambuf> // streambuf
|
|
#include <string> // string, char_traits
|
|
#include <type_traits> // enable_if, is_base_of, is_pointer, is_integral, remove_pointer
|
|
#include <utility> // pair, declval
|
|
|
|
#ifndef JSON_NO_IO
|
|
#include <cstdio> // FILE *
|
|
#include <istream> // istream
|
|
#endif // JSON_NO_IO
|
|
|
|
#include <nlohmann/detail/exceptions.hpp>
|
|
#include <nlohmann/detail/iterators/iterator_traits.hpp>
|
|
#include <nlohmann/detail/macro_scope.hpp>
|
|
#include <nlohmann/detail/meta/type_traits.hpp>
|
|
#include <nlohmann/detail/string_utils.hpp>
|
|
|
|
NLOHMANN_JSON_NAMESPACE_BEGIN
|
|
namespace detail
|
|
{
|
|
|
|
/// the supported input formats
|
|
enum class input_format_t { json, cbor, msgpack, ubjson, bson, bjdata, bon8 };
|
|
|
|
////////////////////
|
|
// input adapters //
|
|
////////////////////
|
|
|
|
#ifndef JSON_NO_IO
|
|
/*!
|
|
Input adapter for stdio file access. This adapter read only 1 byte and do not use any
|
|
buffer. This adapter is a very low level adapter.
|
|
*/
|
|
class file_input_adapter
|
|
{
|
|
public:
|
|
using char_type = char;
|
|
|
|
JSON_HEDLEY_NON_NULL(2)
|
|
explicit file_input_adapter(std::FILE* f) noexcept
|
|
: m_file(f)
|
|
{
|
|
JSON_ASSERT(m_file != nullptr);
|
|
}
|
|
|
|
// make class move-only
|
|
file_input_adapter(const file_input_adapter&) = delete;
|
|
file_input_adapter(file_input_adapter&&) noexcept = default;
|
|
file_input_adapter& operator=(const file_input_adapter&) = delete;
|
|
file_input_adapter& operator=(file_input_adapter&&) = delete;
|
|
~file_input_adapter() = default;
|
|
|
|
std::char_traits<char>::int_type get_character() noexcept
|
|
{
|
|
return std::fgetc(m_file);
|
|
}
|
|
|
|
// returns the number of characters successfully read
|
|
template<class T>
|
|
std::size_t get_elements(T* dest, std::size_t count = 1)
|
|
{
|
|
return fread(dest, 1, sizeof(T) * count, m_file);
|
|
}
|
|
|
|
private:
|
|
/// the file pointer to read from
|
|
std::FILE* m_file;
|
|
};
|
|
|
|
/*!
|
|
Input adapter for a (caching) istream. Does not skip a UTF Byte Order Mark
|
|
itself; that is done by the lexer's skip_bom(). Does not support changing
|
|
the underlying std::streambuf
|
|
in mid-input. Maintains underlying std::istream and std::streambuf to support
|
|
subsequent use of standard std::istream operations to process any input
|
|
characters following those used in parsing the JSON input. Clears the
|
|
std::istream flags; any input errors (e.g., EOF) will be detected by the first
|
|
subsequent call for input from the std::istream.
|
|
*/
|
|
class input_stream_adapter
|
|
{
|
|
public:
|
|
using char_type = char;
|
|
|
|
~input_stream_adapter()
|
|
{
|
|
// clear stream flags; we use underlying streambuf I/O, do not
|
|
// maintain ifstream flags, except eof
|
|
if (is != nullptr)
|
|
{
|
|
#if JSON_PRECISE_STREAM_POSITION
|
|
// consume the character last returned by get_character() unless it
|
|
// was given back with release_lookahead()
|
|
commit_lookahead();
|
|
#endif
|
|
// only call clear() if there is something to clear: it throws
|
|
// std::ios_base::failure if the stream has exceptions() enabled
|
|
// for a state bit that remains set, and a destructor must not throw
|
|
if ((is->rdstate() & ~std::ios::eofbit) != 0)
|
|
{
|
|
is->clear(is->rdstate() & std::ios::eofbit);
|
|
}
|
|
}
|
|
}
|
|
|
|
explicit input_stream_adapter(std::istream& i)
|
|
: is(&i), sb(i.rdbuf())
|
|
{}
|
|
|
|
// deleted because of pointer members
|
|
input_stream_adapter(const input_stream_adapter&) = delete;
|
|
input_stream_adapter& operator=(input_stream_adapter&) = delete;
|
|
input_stream_adapter& operator=(input_stream_adapter&&) = delete;
|
|
|
|
#if JSON_PRECISE_STREAM_POSITION
|
|
input_stream_adapter(input_stream_adapter&& rhs) noexcept
|
|
: is(rhs.is), sb(rhs.sb), lookahead(rhs.lookahead)
|
|
{
|
|
rhs.is = nullptr;
|
|
rhs.sb = nullptr;
|
|
rhs.lookahead = false;
|
|
}
|
|
|
|
// Whether the character last returned by get_character() can be given back
|
|
// to the input with release_lookahead().
|
|
static constexpr bool supports_lookahead = true;
|
|
|
|
// std::istream/std::streambuf use std::char_traits<char>::to_int_type, to
|
|
// ensure that std::char_traits<char>::eof() and the character 0xFF do not
|
|
// end up as the same value, e.g., 0xFFFFFFFF.
|
|
//
|
|
// The character is peeked rather than consumed: it is only stepped over
|
|
// once the next character is requested, or when the adapter is destroyed.
|
|
// Until then, release_lookahead() can leave it in the input.
|
|
std::char_traits<char>::int_type get_character()
|
|
{
|
|
if (lookahead)
|
|
{
|
|
// step over the character returned by the previous call
|
|
sb->sbumpc();
|
|
}
|
|
|
|
auto res = sb->sgetc();
|
|
// set eof manually, as we don't use the istream interface.
|
|
if (JSON_HEDLEY_UNLIKELY(res == std::char_traits<char>::eof()))
|
|
{
|
|
// there is nothing to step over next time
|
|
lookahead = false;
|
|
is->clear(is->rdstate() | std::ios::eofbit);
|
|
}
|
|
else
|
|
{
|
|
lookahead = true;
|
|
}
|
|
return res;
|
|
}
|
|
|
|
// Leave the character last returned by get_character() in the input, so
|
|
// that the next read from the stream - by this adapter or by the caller
|
|
// once parsing is done - sees it again. Unlike putting a consumed
|
|
// character back, this cannot fail.
|
|
void release_lookahead() noexcept
|
|
{
|
|
lookahead = false;
|
|
}
|
|
#else
|
|
input_stream_adapter(input_stream_adapter&& rhs) noexcept
|
|
: is(rhs.is), sb(rhs.sb)
|
|
{
|
|
rhs.is = nullptr;
|
|
rhs.sb = nullptr;
|
|
}
|
|
|
|
// std::istream/std::streambuf use std::char_traits<char>::to_int_type, to
|
|
// ensure that std::char_traits<char>::eof() and the character 0xFF do not
|
|
// end up as the same value, e.g., 0xFFFFFFFF.
|
|
//
|
|
// The character is consumed, so the character that terminates a number
|
|
// stays consumed after parsing; see JSON_PRECISE_STREAM_POSITION.
|
|
std::char_traits<char>::int_type get_character()
|
|
{
|
|
auto res = sb->sbumpc();
|
|
// set eof manually, as we don't use the istream interface.
|
|
if (JSON_HEDLEY_UNLIKELY(res == std::char_traits<char>::eof()))
|
|
{
|
|
is->clear(is->rdstate() | std::ios::eofbit);
|
|
}
|
|
return res;
|
|
}
|
|
#endif
|
|
|
|
template<class T>
|
|
std::size_t get_elements(T* dest, std::size_t count = 1)
|
|
{
|
|
#if JSON_PRECISE_STREAM_POSITION
|
|
commit_lookahead();
|
|
#endif
|
|
auto res = static_cast<std::size_t>(sb->sgetn(reinterpret_cast<char*>(dest), static_cast<std::streamsize>(count * sizeof(T))));
|
|
if (JSON_HEDLEY_UNLIKELY(res < count * sizeof(T)))
|
|
{
|
|
is->clear(is->rdstate() | std::ios::eofbit);
|
|
}
|
|
return res;
|
|
}
|
|
|
|
private:
|
|
#if JSON_PRECISE_STREAM_POSITION
|
|
// Step over the character last returned by get_character(). The character
|
|
// has already been peeked successfully, so for every streambuf with a get
|
|
// area this is a pointer increment that cannot fail.
|
|
void commit_lookahead()
|
|
{
|
|
if (lookahead)
|
|
{
|
|
lookahead = false;
|
|
sb->sbumpc();
|
|
}
|
|
}
|
|
#endif
|
|
|
|
/// the associated input stream
|
|
std::istream* is = nullptr;
|
|
std::streambuf* sb = nullptr;
|
|
#if JSON_PRECISE_STREAM_POSITION
|
|
/// whether get_character() peeked a character that is not consumed yet
|
|
bool lookahead = false;
|
|
#endif
|
|
};
|
|
#endif // JSON_NO_IO
|
|
|
|
// General-purpose iterator-based adapter. It might not be as fast as
|
|
// theoretically possible for some containers, but it is extremely versatile.
|
|
// SentinelType defaults to IteratorType for backward compatibility, but may be
|
|
// a different type, e.g. a C++20 sentinel such as std::default_sentinel_t when
|
|
// IteratorType is a std::counted_iterator.
|
|
template<typename IteratorType, typename SentinelType = IteratorType>
|
|
class iterator_input_adapter
|
|
{
|
|
// Whether the number of elements between two positions can be computed in
|
|
// O(1): either the iterator and the sentinel have the same type (plain
|
|
// std::distance) or, in C++20, the sentinel is a sized sentinel for the
|
|
// iterator (std::ranges::distance), e.g. std::default_sentinel_t paired
|
|
// with std::counted_iterator.
|
|
//
|
|
// JSON_HAS_RANGES gates the C++20 branch: on standard libraries with an
|
|
// incomplete <ranges> (libstdc++ < 11, see #4440) evaluating
|
|
// std::contiguous_iterator on a std::counted_iterator is a hard error
|
|
// instead of yielding false, and these traits are instantiated for every
|
|
// adapter. Such toolchains fall back to the pointer-only test and simply
|
|
// use the byte-at-a-time scanner.
|
|
static constexpr bool sentinel_is_sized =
|
|
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
|
|
std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>;
|
|
#else
|
|
std::is_same<IteratorType, SentinelType>::value;
|
|
#endif
|
|
|
|
public:
|
|
using char_type = typename std::iterator_traits<IteratorType>::value_type;
|
|
|
|
// Whether the lexer may reconstruct already-consumed input on demand (for
|
|
// diagnostics) instead of copying every scanned character eagerly. This is
|
|
// only sound for multi-pass, randomly-addressable byte input: the iterator
|
|
// must be random-access (so the consumed prefix can be revisited in O(1))
|
|
// and each element must map 1:1 to an input byte (wide inputs are wrapped
|
|
// in wide_string_input_adapter, which does not expose this).
|
|
static constexpr bool supports_seek =
|
|
std::is_same<typename std::iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value
|
|
&& sentinel_is_sized
|
|
&& sizeof(char_type) == 1;
|
|
|
|
iterator_input_adapter(IteratorType first, SentinelType last)
|
|
: begin(first), current(std::move(first)), end(std::move(last))
|
|
{}
|
|
|
|
typename char_traits<char_type>::int_type get_character()
|
|
{
|
|
if (JSON_HEDLEY_LIKELY(current != end))
|
|
{
|
|
auto result = char_traits<char_type>::to_int_type(*current);
|
|
std::advance(current, 1);
|
|
return result;
|
|
}
|
|
|
|
return char_traits<char_type>::eof();
|
|
}
|
|
|
|
// number of characters consumed from the input so far
|
|
std::size_t get_consumed_count() const
|
|
{
|
|
return static_cast<std::size_t>(std::distance(begin, current));
|
|
}
|
|
|
|
// append the already-consumed characters in the half-open range
|
|
// [first_index, last_index) to @a out; only valid when supports_seek
|
|
template<typename ContainerType>
|
|
void copy_consumed_range(std::size_t first_index, std::size_t last_index, ContainerType& out) const
|
|
{
|
|
const auto from = std::next(begin, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(first_index));
|
|
const auto to = std::next(begin, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(last_index));
|
|
out.insert(out.end(), from, to);
|
|
}
|
|
|
|
// Copy up to count * sizeof(T) bytes into dest, returning the number of
|
|
// bytes actually read. For contiguous iterators (e.g. pointers) this is a
|
|
// single std::memcpy; for general iterators we fall back to processing the
|
|
// range one-by-one.
|
|
template<class T>
|
|
std::size_t get_elements(T* dest, std::size_t count = 1)
|
|
{
|
|
return get_elements_impl(dest, count, std::integral_constant<bool, iterator_is_contiguous> {});
|
|
}
|
|
|
|
private:
|
|
// whether IteratorType refers to a contiguous range and therefore supports
|
|
// a std::memcpy fast path (pointers always do; in C++20 we can also detect
|
|
// library iterators such as those of std::vector and std::string). The
|
|
// available element count must also be computable in O(1), hence
|
|
// sentinel_is_sized.
|
|
static constexpr bool iterator_is_contiguous = sentinel_is_sized &&
|
|
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
|
|
(std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
|
|
#else
|
|
std::is_pointer<IteratorType>::value;
|
|
#endif
|
|
|
|
// number of unread elements in [current, end)
|
|
std::size_t remaining_count() const
|
|
{
|
|
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
|
|
// std::ranges::distance also supports sized sentinels of a different
|
|
// type (e.g. std::counted_iterator + std::default_sentinel_t)
|
|
return static_cast<std::size_t>(std::ranges::distance(current, end));
|
|
#else
|
|
return static_cast<std::size_t>(std::distance(current, end));
|
|
#endif
|
|
}
|
|
|
|
public:
|
|
// Whether the remaining input is a single contiguous block of 1-byte
|
|
// elements that the lexer can inspect directly (used for the SWAR string
|
|
// fast path).
|
|
static constexpr bool supports_bulk_scan =
|
|
iterator_is_contiguous && sizeof(char_type) == 1;
|
|
|
|
// Pointer to the next unread element; only valid when bulk_remaining() > 0.
|
|
const char_type* bulk_data() const
|
|
{
|
|
return &*current;
|
|
}
|
|
|
|
// Number of unread elements available as one contiguous block.
|
|
std::size_t bulk_remaining() const
|
|
{
|
|
return remaining_count();
|
|
}
|
|
|
|
// Consume @a n elements previously inspected via bulk_data().
|
|
void bulk_skip(std::size_t n)
|
|
{
|
|
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(n));
|
|
}
|
|
|
|
private:
|
|
// contiguous fast path: bulk copy the remaining range with std::memcpy
|
|
template<class T>
|
|
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
|
|
{
|
|
const std::size_t wanted = count * sizeof(T);
|
|
const std::size_t available = remaining_count() * sizeof(char_type);
|
|
const std::size_t copied = (std::min)(wanted, available);
|
|
if (JSON_HEDLEY_LIKELY(copied != 0))
|
|
{
|
|
// the copy must stay within both buffers: the caller-provided
|
|
// destination holds `wanted` bytes and the remaining input range
|
|
// holds `available` bytes, and `copied` is the minimum of the two
|
|
JSON_ASSERT(copied <= wanted); // does not overrun the destination
|
|
JSON_ASSERT(copied <= available); // does not read past the input end
|
|
// &*current yields the raw address for both raw pointers and
|
|
// non-pointer contiguous iterators (e.g. std::vector's iterator)
|
|
std::memcpy(dest, &*current, copied);
|
|
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(copied / sizeof(char_type)));
|
|
}
|
|
return copied;
|
|
}
|
|
|
|
// general fallback: copy the range one element at a time
|
|
template<class T>
|
|
std::size_t get_elements_impl(T* dest, std::size_t count, std::false_type /*contiguous*/)
|
|
{
|
|
auto* ptr = reinterpret_cast<char*>(dest);
|
|
for (std::size_t read_index = 0; read_index < count * sizeof(T); ++read_index)
|
|
{
|
|
if (JSON_HEDLEY_LIKELY(current != end))
|
|
{
|
|
ptr[read_index] = static_cast<char>(*current);
|
|
std::advance(current, 1);
|
|
}
|
|
else
|
|
{
|
|
return read_index;
|
|
}
|
|
}
|
|
return count * sizeof(T);
|
|
}
|
|
|
|
IteratorType begin;
|
|
IteratorType current;
|
|
SentinelType end;
|
|
|
|
template<typename BaseInputAdapter, size_t T>
|
|
friend struct wide_string_input_helper;
|
|
|
|
bool empty() const
|
|
{
|
|
return current == end;
|
|
}
|
|
};
|
|
|
|
template<typename BaseInputAdapter, size_t T>
|
|
struct wide_string_input_helper;
|
|
|
|
template<typename BaseInputAdapter>
|
|
struct wide_string_input_helper<BaseInputAdapter, 4>
|
|
{
|
|
// UTF-32
|
|
static void fill_buffer(BaseInputAdapter& input,
|
|
std::array<std::char_traits<char>::int_type, 4>& utf8_bytes,
|
|
size_t& utf8_bytes_index,
|
|
size_t& utf8_bytes_filled)
|
|
{
|
|
utf8_bytes_index = 0;
|
|
|
|
if (JSON_HEDLEY_UNLIKELY(input.empty()))
|
|
{
|
|
utf8_bytes[0] = std::char_traits<char>::eof();
|
|
utf8_bytes_filled = 1;
|
|
}
|
|
else
|
|
{
|
|
// get the current character
|
|
const auto wc = input.get_character();
|
|
|
|
if (wc <= 0x10FFFF)
|
|
{
|
|
// UTF-32 to UTF-8 encoding
|
|
utf8_bytes_filled = 0;
|
|
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
|
{
|
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
|
});
|
|
}
|
|
else
|
|
{
|
|
// A code point above U+10FFFF has no UTF-8 encoding. Passing the
|
|
// unit through would narrow it to int, where 0xFFFFFFFF becomes
|
|
// char_traits<char>::eof() and would end the input silently, so
|
|
// emit a byte that is never valid UTF-8 and let the decoder
|
|
// reject it.
|
|
utf8_bytes[0] = 0xFF;
|
|
utf8_bytes_filled = 1;
|
|
}
|
|
}
|
|
}
|
|
};
|
|
|
|
template<typename BaseInputAdapter>
|
|
struct wide_string_input_helper<BaseInputAdapter, 2>
|
|
{
|
|
// UTF-16
|
|
static void fill_buffer(BaseInputAdapter& input,
|
|
std::array<std::char_traits<char>::int_type, 4>& utf8_bytes,
|
|
size_t& utf8_bytes_index,
|
|
size_t& utf8_bytes_filled)
|
|
{
|
|
utf8_bytes_index = 0;
|
|
|
|
if (JSON_HEDLEY_UNLIKELY(input.empty()))
|
|
{
|
|
utf8_bytes[0] = std::char_traits<char>::eof();
|
|
utf8_bytes_filled = 1;
|
|
}
|
|
else
|
|
{
|
|
// get the current character
|
|
const auto wc = input.get_character();
|
|
|
|
if (0xD800 > wc || wc >= 0xE000)
|
|
{
|
|
// a UTF-16 code unit outside the surrogate range is a valid
|
|
// code point (at most U+FFFF) on its own
|
|
utf8_bytes_filled = 0;
|
|
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
|
{
|
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
|
});
|
|
}
|
|
else
|
|
{
|
|
// A supplementary code point is a high surrogate (0xD800..0xDBFF)
|
|
// followed by a low surrogate (0xDC00..0xDFFF). A lone low
|
|
// surrogate, a high surrogate at the end of the input, or a high
|
|
// surrogate followed by any other unit is malformed UTF-16. In
|
|
// that case the offending unit is passed through unchanged so the
|
|
// UTF-8 decoder rejects it, matching how \uXXXX surrogate escapes
|
|
// are handled in the lexer.
|
|
bool valid_pair = false;
|
|
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
|
|
{
|
|
const auto wc2 = static_cast<unsigned int>(input.get_character());
|
|
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
|
|
{
|
|
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
|
|
utf8_bytes_filled = 0;
|
|
encode_utf8(charcode, [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
|
{
|
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
|
});
|
|
valid_pair = true;
|
|
}
|
|
}
|
|
|
|
if (!valid_pair)
|
|
{
|
|
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
|
|
utf8_bytes_filled = 1;
|
|
}
|
|
}
|
|
}
|
|
}
|
|
};
|
|
|
|
// Wraps another input adapter to convert wide character types into individual bytes.
|
|
template<typename BaseInputAdapter, typename WideCharType>
|
|
class wide_string_input_adapter
|
|
{
|
|
public:
|
|
using char_type = char;
|
|
|
|
wide_string_input_adapter(BaseInputAdapter base)
|
|
: base_adapter(base) {}
|
|
|
|
typename std::char_traits<char>::int_type get_character() noexcept
|
|
{
|
|
// check if the buffer needs to be filled
|
|
if (utf8_bytes_index == utf8_bytes_filled)
|
|
{
|
|
fill_buffer<sizeof(WideCharType)>();
|
|
|
|
JSON_ASSERT(utf8_bytes_filled > 0);
|
|
JSON_ASSERT(utf8_bytes_index == 0);
|
|
}
|
|
|
|
// use buffer
|
|
JSON_ASSERT(utf8_bytes_filled > 0);
|
|
JSON_ASSERT(utf8_bytes_index < utf8_bytes_filled);
|
|
return utf8_bytes[utf8_bytes_index++];
|
|
}
|
|
|
|
// parsing binary with wchar doesn't make sense, but since the parsing mode can be runtime, we need something here
|
|
template<class T>
|
|
JSON_HEDLEY_NO_RETURN std::size_t get_elements(T* /*dest*/, std::size_t /*count*/ = 1)
|
|
{
|
|
JSON_THROW(parse_error::create(112, 1, "wide string type cannot be interpreted as binary data", nullptr));
|
|
}
|
|
|
|
private:
|
|
BaseInputAdapter base_adapter;
|
|
|
|
template<size_t T>
|
|
void fill_buffer()
|
|
{
|
|
wide_string_input_helper<BaseInputAdapter, T>::fill_buffer(base_adapter, utf8_bytes, utf8_bytes_index, utf8_bytes_filled);
|
|
}
|
|
|
|
/// a buffer for UTF-8 bytes
|
|
std::array<std::char_traits<char>::int_type, 4> utf8_bytes = {{0, 0, 0, 0}};
|
|
|
|
/// index to the utf8_codes array for the next valid byte
|
|
std::size_t utf8_bytes_index = 0;
|
|
/// number of valid bytes in the utf8_codes array
|
|
std::size_t utf8_bytes_filled = 0;
|
|
};
|
|
|
|
template<typename IteratorType, typename SentinelType = IteratorType, typename Enable = void>
|
|
struct iterator_input_adapter_factory
|
|
{
|
|
using iterator_type = IteratorType;
|
|
using sentinel_type = SentinelType;
|
|
using char_type = typename std::iterator_traits<iterator_type>::value_type;
|
|
using adapter_type = iterator_input_adapter<iterator_type, sentinel_type>;
|
|
|
|
static adapter_type create(IteratorType first, SentinelType last)
|
|
{
|
|
return adapter_type(std::move(first), std::move(last));
|
|
}
|
|
};
|
|
|
|
// Detection: whether IteratorType and SentinelType can be compared with !=
|
|
template<typename IteratorType, typename SentinelType, typename = void>
|
|
struct can_compare_ne_impl : std::false_type {};
|
|
|
|
template<typename IteratorType, typename SentinelType>
|
|
struct can_compare_ne_impl < IteratorType, SentinelType,
|
|
void_t < decltype(std::declval<IteratorType>() != std::declval<SentinelType>()) >>
|
|
: std::true_type {};
|
|
|
|
// Workaround for reversed operator order
|
|
template<typename IteratorType, typename SentinelType, typename = void>
|
|
struct can_compare_ne_reversed : std::false_type {};
|
|
|
|
template<typename IteratorType, typename SentinelType>
|
|
struct can_compare_ne_reversed < IteratorType, SentinelType,
|
|
void_t < decltype(std::declval<SentinelType>() != std::declval<IteratorType>()) >>
|
|
: std::true_type {};
|
|
|
|
template<typename IteratorType, typename SentinelType>
|
|
struct can_compare_ne_either_order : std::integral_constant < bool,
|
|
can_compare_ne_impl<IteratorType, SentinelType>::value ||
|
|
can_compare_ne_reversed<IteratorType, SentinelType>::value > {};
|
|
|
|
// std::nullptr_t is excluded explicitly: a literal `nullptr` passed as a
|
|
// trailing default argument (e.g. parse(s, nullptr, ...)) must never be
|
|
// mistaken for a sentinel, and some compilers (e.g. GCC 4.8) unreliably
|
|
// SFINAE the `operator!=` detection above for std::nullptr_t against
|
|
// container/string types, which would otherwise make such calls ambiguous
|
|
// with the compatible-input overload.
|
|
template<typename IteratorType, typename SentinelType>
|
|
struct can_compare_ne : std::integral_constant < bool,
|
|
!std::is_same<SentinelType, std::nullptr_t>::value &&
|
|
can_compare_ne_either_order<IteratorType, SentinelType>::value > {};
|
|
|
|
template<typename T>
|
|
struct is_iterator_of_multibyte
|
|
{
|
|
using value_type = typename std::iterator_traits<T>::value_type;
|
|
enum // NOLINT(cppcoreguidelines-use-enum-class)
|
|
{
|
|
value = sizeof(value_type) > 1
|
|
};
|
|
};
|
|
|
|
template<typename IteratorType, typename SentinelType>
|
|
struct iterator_input_adapter_factory<IteratorType, SentinelType, enable_if_t<is_iterator_of_multibyte<IteratorType>::value>>
|
|
{
|
|
using iterator_type = IteratorType;
|
|
using sentinel_type = SentinelType;
|
|
using char_type = typename std::iterator_traits<iterator_type>::value_type;
|
|
using base_adapter_type = iterator_input_adapter<iterator_type, sentinel_type>;
|
|
using adapter_type = wide_string_input_adapter<base_adapter_type, char_type>;
|
|
|
|
static adapter_type create(IteratorType first, SentinelType last)
|
|
{
|
|
return adapter_type(base_adapter_type(std::move(first), std::move(last)));
|
|
}
|
|
};
|
|
|
|
// General purpose iterator-based input (iterator+sentinel pair; SentinelType
|
|
// defaults to IteratorType for the common same-type case, but may differ for
|
|
// C++20 ranges-style iterator+sentinel pairs). Only enable for types that can
|
|
// be compared with !=.
|
|
template < typename IteratorType, typename SentinelType = IteratorType,
|
|
typename = typename std::enable_if <
|
|
can_compare_ne<IteratorType, SentinelType>::value >::type >
|
|
typename iterator_input_adapter_factory<IteratorType, SentinelType>::adapter_type input_adapter(IteratorType first, SentinelType last)
|
|
{
|
|
using factory_type = iterator_input_adapter_factory<IteratorType, SentinelType>;
|
|
return factory_type::create(first, last);
|
|
}
|
|
|
|
// The element type a container's data() points at, cv-qualifiers removed.
|
|
// Ill-formed - and therefore SFINAE-friendly - for types without data().
|
|
template<typename ContainerType>
|
|
using container_data_t = typename std::remove_cv<typename std::remove_pointer <
|
|
decltype(std::declval<const ContainerType&>().data()) >::type >::type;
|
|
|
|
// The container's own element type, cv-qualifiers removed. It is looked up on
|
|
// the bare type so it is also found when ContainerType is deduced as a
|
|
// reference by the forwarding-reference overload below.
|
|
template<typename ContainerType>
|
|
using container_value_t = typename std::remove_cv <
|
|
typename std::remove_cv<typename std::remove_reference<ContainerType>::type>::type::value_type >::type;
|
|
|
|
// Detect a container that stores its elements contiguously as single bytes
|
|
// (std::string, std::vector<char/unsigned char>, std::array<char, N>,
|
|
// std::string_view, ...). Such inputs are wrapped in a pointer-based adapter so
|
|
// they benefit from the contiguous fast paths (bulk string scanning, memcpy for
|
|
// binary formats) in every C++ standard - not only in C++20, where the standard
|
|
// library iterators model std::contiguous_iterator and are detected directly.
|
|
//
|
|
// data() and size() on their own would be duck typing: they say nothing about
|
|
// size() counting the units data() points at, and reading [data(), data() +
|
|
// size()) as bytes would be wrong for a type where it does not. Requiring the
|
|
// container's own value_type to be that same single-byte element ties the two
|
|
// together; every contiguous standard container satisfies it. Anything else
|
|
// keeps the iterator-based adapter, which is always correct - only slower.
|
|
template<typename ContainerType, typename = void>
|
|
struct is_contiguous_byte_container : std::false_type {};
|
|
|
|
template<typename ContainerType>
|
|
struct is_contiguous_byte_container < ContainerType, void_t <
|
|
container_data_t<ContainerType>,
|
|
container_value_t<ContainerType>,
|
|
decltype(std::declval<const ContainerType&>().size()) >>
|
|
: std::integral_constant < bool,
|
|
std::is_pointer<decltype(std::declval<const ContainerType&>().data())>::value&&
|
|
std::is_integral<container_data_t<ContainerType>>::value&&
|
|
sizeof(container_data_t<ContainerType>) == 1 &&
|
|
std::is_same<container_data_t<ContainerType>, container_value_t<ContainerType>>::value > {};
|
|
|
|
// Convenience shorthand from container to iterator
|
|
// Enables ADL on begin(container) and end(container)
|
|
// Encloses the using declarations in namespace for not to leak them to outside scope
|
|
|
|
namespace container_input_adapter_factory_impl
|
|
{
|
|
|
|
using std::begin;
|
|
using std::end;
|
|
|
|
template<typename ContainerType, typename Enable = void>
|
|
struct container_input_adapter_factory {};
|
|
|
|
template<typename ContainerType>
|
|
struct container_input_adapter_factory< ContainerType,
|
|
void_t<decltype(begin(std::declval<ContainerType>()), end(std::declval<ContainerType>()))>>
|
|
{
|
|
using adapter_type = decltype(input_adapter(begin(std::declval<ContainerType>()), end(std::declval<ContainerType>())));
|
|
|
|
static adapter_type create(ContainerType&& container)
|
|
{
|
|
// container is forwarded twice on purpose: the resulting begin/end
|
|
// iterator types must match adapter_type, computed the same way
|
|
// NOLINTNEXTLINE(bugprone-use-after-move)
|
|
return input_adapter(begin(std::forward<ContainerType>(container)), end(std::forward<ContainerType>(container)));
|
|
}
|
|
};
|
|
|
|
} // namespace container_input_adapter_factory_impl
|
|
|
|
// General container path (iterator-based). Contiguous single-byte containers
|
|
// are excluded here and routed through the pointer-based overload below.
|
|
template < typename ContainerType,
|
|
enable_if_t < !is_contiguous_byte_container<ContainerType>::value, int > = 0 >
|
|
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType && container)
|
|
{
|
|
return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container));
|
|
}
|
|
|
|
// Contiguous single-byte containers (std::string, std::vector<char>, ...) are
|
|
// wrapped in a pointer-based adapter so the contiguous fast paths apply in every
|
|
// standard. The pointer keeps the container's own element type (const char* for
|
|
// std::string, const std::uint8_t* for std::vector<std::uint8_t>, ...), so the
|
|
// resulting char_type - and therefore the parsing behavior - is byte-for-byte
|
|
// identical to the iterator-based path; only the raw pointer additionally
|
|
// enables the bulk fast paths. The container outlives the adapter for the whole
|
|
// parse (temporaries live until the end of the full expression), exactly as the
|
|
// iterators it replaces did.
|
|
template < typename ContainerType,
|
|
enable_if_t < is_contiguous_byte_container<ContainerType>::value, int > = 0 >
|
|
auto input_adapter(const ContainerType& container)
|
|
-> decltype(input_adapter(container.data(), container.data() + container.size()))
|
|
{
|
|
return input_adapter(container.data(), container.data() + container.size());
|
|
}
|
|
|
|
// specialization for std::string
|
|
using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>()));
|
|
|
|
#ifndef JSON_NO_IO
|
|
// Special cases with fast paths
|
|
inline file_input_adapter input_adapter(std::FILE* file)
|
|
{
|
|
if (file == nullptr)
|
|
{
|
|
JSON_THROW(parse_error::create(101, 0, "attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
|
|
}
|
|
return file_input_adapter(file);
|
|
}
|
|
|
|
inline input_stream_adapter input_adapter(std::istream& stream)
|
|
{
|
|
if (stream.rdbuf() == nullptr)
|
|
{
|
|
JSON_THROW(parse_error::create(101, 0, "attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
|
|
}
|
|
return input_stream_adapter(stream);
|
|
}
|
|
|
|
inline input_stream_adapter input_adapter(std::istream&& stream)
|
|
{
|
|
return input_adapter(stream);
|
|
}
|
|
#endif // JSON_NO_IO
|
|
|
|
using contiguous_bytes_input_adapter = decltype(input_adapter(std::declval<const char*>(), std::declval<const char*>()));
|
|
|
|
// Null-delimited strings, and the like.
|
|
template < typename CharT,
|
|
typename std::enable_if <
|
|
std::is_pointer<CharT>::value&&
|
|
!std::is_array<CharT>::value&&
|
|
std::is_integral<typename std::remove_pointer<CharT>::type>::value&&
|
|
sizeof(typename std::remove_pointer<CharT>::type) == 1,
|
|
int >::type = 0 >
|
|
contiguous_bytes_input_adapter input_adapter(CharT b)
|
|
{
|
|
if (b == nullptr)
|
|
{
|
|
JSON_THROW(parse_error::create(101, 0, "attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
|
|
}
|
|
auto length = std::strlen(reinterpret_cast<const char*>(b));
|
|
const auto* ptr = reinterpret_cast<const char*>(b);
|
|
return input_adapter(ptr, ptr + length); // cppcheck-suppress[nullPointerArithmeticRedundantCheck]
|
|
}
|
|
|
|
template<typename T, std::size_t N>
|
|
auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
|
|
{
|
|
#if JSON_STRICT_NUL_HANDLING
|
|
// A text-literal array from string-literal initialization (e.g.
|
|
// json::parse("123") or json::parse(L"123")) carries a trailing '\0'
|
|
// contributed by the compiler, not by the source text; drop exactly that
|
|
// one byte so it is not mistaken for real trailing data. This covers all
|
|
// character types that string literals can use: char, wchar_t, char16_t,
|
|
// char32_t, and (C++20) char8_t. Every other element type (unsigned char,
|
|
// std::uint8_t, ...) keeps the full extent unconditionally, since a
|
|
// trailing zero byte there is genuine data (e.g. CBOR/MessagePack). This
|
|
// intentionally does not strlen()-scan the array (as the pointer overload
|
|
// above does for a null-delimited string): for an array that is not
|
|
// NUL-terminated within its bounds, that would read past the end of the
|
|
// array.
|
|
using char_t = typename std::remove_cv<T>::type;
|
|
constexpr bool is_text_literal_type = std::is_same<char_t, char>::value
|
|
|| std::is_same<char_t, wchar_t>::value
|
|
|| std::is_same<char_t, char16_t>::value
|
|
|| std::is_same<char_t, char32_t>::value
|
|
#if defined(__cpp_char8_t)
|
|
|| std::is_same<char_t, char8_t>::value
|
|
#endif
|
|
;
|
|
if (is_text_literal_type && N > 0 && array[N - 1] == 0)
|
|
{
|
|
return input_adapter(array, array + N - 1);
|
|
}
|
|
#endif
|
|
return input_adapter(array, array + N);
|
|
}
|
|
|
|
// This class only handles inputs that construct a contiguous_bytes_input_adapter
|
|
// (e.g. span_input_adapter). It's required so that expressions like {ptr, len}
|
|
// can be implicitly cast to the correct adapter.
|
|
class span_input_adapter
|
|
{
|
|
public:
|
|
template < typename CharT,
|
|
typename std::enable_if <
|
|
std::is_pointer<CharT>::value&&
|
|
std::is_integral<typename std::remove_pointer<CharT>::type>::value&&
|
|
sizeof(typename std::remove_pointer<CharT>::type) == 1,
|
|
int >::type = 0 >
|
|
span_input_adapter(CharT b, std::size_t l)
|
|
: ia(reinterpret_cast<const char*>(b), reinterpret_cast<const char*>(b) + l) {}
|
|
|
|
template<class IteratorType,
|
|
typename std::enable_if<
|
|
std::is_same<typename iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value,
|
|
int>::type = 0>
|
|
span_input_adapter(IteratorType first, IteratorType last)
|
|
: ia(input_adapter(first, last)) {}
|
|
|
|
contiguous_bytes_input_adapter&& get()
|
|
{
|
|
return std::move(ia); // NOLINT(hicpp-move-const-arg,performance-move-const-arg)
|
|
}
|
|
|
|
private:
|
|
contiguous_bytes_input_adapter ia;
|
|
};
|
|
|
|
} // namespace detail
|
|
NLOHMANN_JSON_NAMESPACE_END
|