Scan the strings of json_view with NEON and SSE2

Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.

Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.

json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.

Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-28 23:35:37 +02:00
parent a491cc7663
commit ebfbbf29cb
15 changed files with 863 additions and 20 deletions

View File

@@ -136,6 +136,32 @@ void check_same(const std::string& text)
}
}
// a string value and a key must be accepted or rejected as json::parse does,
// and give its value (cheaper than check_same: the options do not matter)
void check_string(const std::string& content)
{
for (const std::string& text :
{
"[\"" + content + "\"]", "{\"" + content + "\":1}"
})
{
const bool accepted = json::accept(text);
for (const bool sentinel :
{
true, false
})
{
const built b = build(text, false, false, sentinel);
if (b.ok != accepted || (accepted && value_of(b) != json::parse(text)))
{
CAPTURE(text);
CHECK(b.ok == accepted);
CHECK((b.ok && accepted ? value_of(b) == json::parse(text) : true));
}
}
}
}
// a small deterministic generator of documents
struct generator
{
@@ -364,3 +390,81 @@ TEST_CASE("json_view builder")
}
}
}
TEST_CASE("json_view builder: strings across vector blocks")
{
// Strings are scanned 8 or 16 bytes at a time (NEON, SSE2, or SWAR) and
// non-ASCII text with the vector UTF-8 check (NEON, SSSE3) or one sequence
// at a time. Sequences are placed so that they start at every offset
// around the block boundaries of keys (16, 32) and values (8, 24), with
// text of several lengths after them.
const std::size_t prefixes[] = {0, 6, 7, 8, 13, 14, 15, 16, 21, 22, 23, 24, 29, 30, 31, 32};
const std::size_t suffixes[] = {0, 3, 17};
const auto around = [&](const std::string & seq, std::size_t prefix, std::size_t suffix)
{
return std::string(prefix, 'a') + seq + std::string(suffix, 'b');
};
SECTION("every two-byte sequence")
{
for (unsigned lead = 0x80; lead <= 0xFF; ++lead)
{
for (unsigned second = 0; second <= 0xFF; ++second)
{
if (second == '"' || second == '\\')
{
continue;
}
const std::string seq = {static_cast<char>(lead), static_cast<char>(second)};
check_string(around(seq, prefixes[(lead + second) % 16], suffixes[second % 3]));
}
}
}
SECTION("three- and four-byte sequences")
{
const unsigned conts[] = {0x7F, 0x80, 0x8F, 0x90, 0x9F, 0xA0, 0xBF, 0xC0};
for (unsigned lead = 0xE0; lead <= 0xF7; ++lead)
{
for (const unsigned b2 : conts)
{
for (const unsigned b3 : conts)
{
for (const std::size_t prefix : prefixes)
{
std::string seq = {static_cast<char>(lead), static_cast<char>(b2), static_cast<char>(b3)};
if (lead >= 0xF0)
{
seq += static_cast<char>(prefix % 2 == 0 ? 0x80 : 0xBF);
}
check_string(around(seq, prefix, suffixes[prefix % 3]));
// cut short before the end of the string
check_string(around(seq.substr(0, seq.size() - 1), prefix, suffixes[prefix % 3]));
}
}
}
}
}
SECTION("long runs of text with one damaged byte")
{
const char* const chars[] = {"a", "\xc3\xa9", "\xe3\x81\x82", "\xf0\x9f\x98\x80", "\xed\x9f\xbf", "\xef\xbf\xbf", "\xf4\x8f\xbf\xbf"};
const char damage[] = {'\x80', '\xbf', '\xc0', '\xc1', '\xe0', '\xed', '\xf5', '\xff', '\x1f', '"', '\\'};
std::mt19937 rng(5295);
for (int i = 0; i < 4000; ++i)
{
std::string text;
const auto n = rng() % 60;
for (unsigned k = 0; k < n; ++k)
{
text += chars[rng() % 7];
}
check_string(text);
if (!text.empty())
{
text[rng() % text.size()] = damage[rng() % 11];
check_string(text);
}
}
}
}