From c1bcda8dd11433e09bf7faf8f6722af50c36612a Mon Sep 17 00:00:00 2001 From: Aaron Kimbrell Date: Wed, 30 Sep 2026 08:05:00 -0500 Subject: [PATCH 1/4] fix(chat-filter): portable .dcf hashing, block list phrases The filter stored and compared words by std::hash in size_t, which differs between standard libraries and platforms, so a .dcf made on one system never matched on another and the block list never worked there (issue 215). - ChatFilterCore.h: 64-bit FNV-1a over the entry's bytes, ASCII lower case, fixed-width uint64_t everywhere stored or compared. - .dcf version 3: little-endian magic, version, longest entry in words, uint64 count and sorted uint64 hashes. Version 2 files are refused: the allowed words cache is rebuilt from its .txt, an old blocklist.dcf is logged as unreadable. - The servers build blocklist.dcf from a plain blocklist.txt next to them (one word or phrase per line) when it is newer. - Blocked entries can be phrases: runs of consecutive words up to the longest entry, the whole run marked. Whitelist chat still checks one word at a time, as the client does. - Dashboard: the chat filter API reads blocklist.dcf the same way (status, phrase length), accepts blocked phrases, refuses allowed ones, and explains phrase matches in its message test. Co-Authored-By: Claude Opus 5.5 --- dChatFilter/ChatFilterCore.h | 342 ++++++++++++++++++++ dChatFilter/dChatFilter.cpp | 209 +++++------- dChatFilter/dChatFilter.h | 54 ++-- dDashboardServer/routes/ChatFilterWords.h | 97 +++--- dDashboardServer/routes/ModerationTools.cpp | 55 ++-- tests/dCommonTests/CMakeLists.txt | 3 + tests/dCommonTests/ChatFilterCoreTests.cpp | 149 +++++++++ tests/dWebTests/ModerationToolsTests.cpp | 50 +-- 8 files changed, 691 insertions(+), 268 deletions(-) create mode 100644 dChatFilter/ChatFilterCore.h create mode 100644 tests/dCommonTests/ChatFilterCoreTests.cpp diff --git a/dChatFilter/ChatFilterCore.h b/dChatFilter/ChatFilterCore.h new file mode 100644 index 000000000..b6acbb609 --- /dev/null +++ b/dChatFilter/ChatFilterCore.h @@ -0,0 +1,342 @@ +#pragma once + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +/** + * The chat filter's words, without the server around them (pure, unit tested; dChatFilter and the dashboard both use it). + * + * A word is compared lower case (ASCII only, so every platform agrees) without ! ? ; . , and a phrase is its words + * joined by one space. Entries are stored and compared by ChatFilterWords::Hash: 64-bit FNV-1a over the entry's bytes, + * the same on every compiler, standard library and platform. + */ +namespace ChatFilterWords { + // ASCII lower case; other bytes (UTF-8) stay as they are + inline std::string AsciiLower(std::string text) { + for (auto& c : text) if (c >= 'A' && c <= 'Z') c = static_cast(c - 'A' + 'a'); + return text; + } + + // A word as the filter compares it: lower case, without ! ? ; . , + inline std::string NormalizeWord(std::string word) { + std::erase_if(word, [](char c) { return c == '!' || c == '?' || c == ';' || c == '.' || c == ','; }); + return AsciiLower(std::move(word)); + } + + // A word or phrase as the filter stores it: each word normalized, words that end up empty dropped, joined by one space + inline std::string NormalizeEntry(std::string_view text) { + std::string entry; + size_t start = 0; + while (start < text.size()) { + auto end = text.find_first_of(" \t\r\n", start); + if (end == std::string_view::npos) end = text.size(); + const auto word = NormalizeWord(std::string(text.substr(start, end - start))); + if (!word.empty()) { + if (!entry.empty()) entry += ' '; + entry += word; + } + start = end + 1; + } + return entry; + } + + // How many words an entry has (1 for a word, 0 for an empty entry) + inline uint32_t WordCount(std::string_view entry) { + return entry.empty() ? 0 : static_cast(std::count(entry.begin(), entry.end(), ' ')) + 1; + } + + // 64-bit FNV-1a: offset basis 0xcbf29ce484222325, prime 0x100000001b3, one byte at a time + constexpr uint64_t Hash(std::string_view entry) { + uint64_t hash = 0xcbf29ce484222325ULL; + for (const char c : entry) { + hash ^= static_cast(c); + hash *= 0x100000001b3ULL; + } + return hash; + } + + // One piece of a message between spaces: where it is in the message and the word the filter compares + struct Token { + uint32_t position{}; + uint32_t length{}; + std::string word; + }; + + // A message split at each space, the way the filter checks it: two spaces in a row give an empty piece, a trailing space none + inline std::vector Tokenize(std::string_view message) { + std::vector tokens; + size_t start = 0; + while (start < message.size()) { + auto end = message.find(' ', start); + if (end == std::string_view::npos) end = message.size(); + tokens.push_back({ static_cast(start), static_cast(end - start), NormalizeWord(std::string(message.substr(start, end - start))) }); + start = end + 1; + } + return tokens; + } + + // A run of tokens [first, last] that matched a blocked entry, and the longest entry that matched where the run starts + struct Match { + size_t first{}; + size_t last{}; + std::string entry; + }; + + /** + * The blocked words and phrases in a message: at each word, the longest run of up to maxWords consecutive words + * (tokens with no word are skipped) whose entry isBlocked accepts. Runs that share a word are merged into one. + */ + inline std::vector FindBlocked(const std::vector& tokens, uint32_t maxWords, const std::function& isBlocked) { + std::vector words; + for (size_t i = 0; i < tokens.size(); i++) if (!tokens[i].word.empty()) words.push_back(i); + + std::vector matches; + for (size_t i = 0; i < words.size(); i++) { + std::string entry; + std::string best; + size_t bestLength = 0; + for (size_t n = 1; n <= maxWords && i + n <= words.size(); n++) { + if (n > 1) entry += ' '; + entry += tokens[words[i + n - 1]].word; + if (isBlocked(entry)) { + best = entry; + bestLength = n; + } + } + if (bestLength == 0) continue; + const size_t first = words[i]; + const size_t last = words[i + bestLength - 1]; + if (!matches.empty() && first <= matches.back().last) { + matches.back().last = std::max(matches.back().last, last); + } else { + matches.push_back({ first, last, std::move(best) }); + } + } + return matches; + } + + // Hashes of words or phrases, and the most words any of them has + struct WordList { + std::unordered_set hashes; + uint32_t maxWords{}; + + void AddEntry(std::string_view entry) { + if (entry.empty()) return; + hashes.insert(Hash(entry)); + maxWords = std::max(maxWords, WordCount(entry)); + } + bool Contains(std::string_view entry) const { return hashes.contains(Hash(entry)); } + bool Empty() const { return hashes.empty(); } + size_t Size() const { return hashes.size(); } + }; + + // Everything the filter checks a message against + struct Lists { + WordList approved; // chatplus_en_us.txt and approved character names: whitelist chat, one word at a time + WordList denied; // blocklist.dcf: best friends' free chat + WordList customAllowed; // allowed on the dashboard + WordList customBlocked; // blocked on the dashboard: stopped in every kind of chat + }; + + /** + * The pieces of a message the filter stops, as (position, length) in the message. Blocked words and phrases (the + * dashboard's always, blocklist.dcf's in free chat) are stopped as one span each. In whitelist chat (allowList) every + * other piece must be an allowed word, one at a time, as the client checks words. In free chat without a block list + * the whole message is stopped. + */ + inline std::set> CheckMessage(std::string_view message, bool allowList, const Lists& lists) { + if (message.empty()) return {}; + if (!allowList && lists.denied.Empty()) return { { 0, static_cast(message.length()) } }; + + const auto tokens = Tokenize(message); + const uint32_t maxWords = std::max(lists.customBlocked.maxWords, allowList ? 0u : lists.denied.maxWords); + const auto matches = FindBlocked(tokens, maxWords, [&](const std::string& entry) { + return lists.customBlocked.Contains(entry) || (!allowList && lists.denied.Contains(entry)); + }); + + std::set> bad; + std::vector covered(tokens.size(), false); + for (const auto& match : matches) { + const auto& first = tokens[match.first]; + const auto& last = tokens[match.last]; + bad.emplace(static_cast(first.position), static_cast(last.position + last.length - first.position)); + for (size_t i = match.first; i <= match.last; i++) covered[i] = true; + } + + if (allowList) { + for (size_t i = 0; i < tokens.size(); i++) { + if (covered[i]) continue; + const auto hash = Hash(tokens[i].word); + if (!lists.approved.hashes.contains(hash) && !lists.customAllowed.hashes.contains(hash)) { + bad.emplace(static_cast(tokens[i].position), static_cast(tokens[i].length)); + } + } + } + return bad; + } +} + +/** + * The chat filter's word list files (.dcf). These are DLU's own files: the client reads no .dcf and hashes no chat + * words (it keeps its lists as plain text). Layout, little-endian: + * uint32 magic 'DCFB' | uint32 version (3) | uint32 most words in one entry | uint64 count | count x uint64 ChatFilterWords::Hash + * Version 2 (older DLU) stored std::hash values, which differ between compilers and platforms; those can't be read. + */ +namespace dChatFilterDCF { + constexpr uint32_t header = ('D' + ('C' << 8) + ('F' << 16) + ('B' << 24)); + constexpr uint32_t formatVersion = 3; + constexpr uint32_t oldFormatVersion = 2; + constexpr size_t headerSize = 4 + 4 + 4 + 8; + + // The block list's plain source (one word or phrase per line) and the .dcf built from it, next to the servers + constexpr const char* BLOCK_LIST_TEXT = "blocklist.txt"; + constexpr const char* BLOCK_LIST_FILE = "blocklist.dcf"; + + enum class eStatus : uint8_t { + OK, + MISSING, // no file + NOT_DCF, // not a .dcf file + OLD_FORMAT, // version 2: platform-dependent hashes, rebuild it from the plain word list + UNKNOWN, // a version this server doesn't know + TRUNCATED, // shorter than its count says + }; + + inline const char* StatusText(eStatus status) { + switch (status) { + case eStatus::OK: return "ok"; + case eStatus::MISSING: return "missing"; + case eStatus::NOT_DCF: return "not a .dcf file"; + case eStatus::OLD_FORMAT: return "old format (version 2, platform-dependent hashes)"; + case eStatus::UNKNOWN: return "unknown version"; + case eStatus::TRUNCATED: return "truncated"; + } + return "unknown"; + } + + struct ParseResult { + eStatus status{ eStatus::NOT_DCF }; + uint32_t version{}; + ChatFilterWords::WordList list; + }; + + namespace detail { + inline uint64_t ReadLE(std::string_view bytes, size_t offset, size_t size) { + uint64_t value = 0; + for (size_t i = 0; i < size; i++) value |= static_cast(static_cast(bytes[offset + i])) << (8 * i); + return value; + } + inline void WriteLE(std::string& out, uint64_t value, size_t size) { + for (size_t i = 0; i < size; i++) out.push_back(static_cast((value >> (8 * i)) & 0xFF)); + } + } + + inline ParseResult Parse(std::string_view bytes) { + ParseResult result; + if (bytes.size() < 8 || detail::ReadLE(bytes, 0, 4) != header) return result; + result.version = static_cast(detail::ReadLE(bytes, 4, 4)); + if (result.version == oldFormatVersion) { + result.status = eStatus::OLD_FORMAT; + return result; + } + if (result.version != formatVersion) { + result.status = eStatus::UNKNOWN; + return result; + } + result.status = eStatus::TRUNCATED; + if (bytes.size() < headerSize) return result; + const auto maxWords = static_cast(detail::ReadLE(bytes, 8, 4)); + const auto count = detail::ReadLE(bytes, 12, 8); + if (count > (bytes.size() - headerSize) / 8) return result; + result.list.maxWords = maxWords; + result.list.hashes.reserve(count); + for (uint64_t i = 0; i < count; i++) result.list.hashes.insert(detail::ReadLE(bytes, headerSize + i * 8, 8)); + result.status = eStatus::OK; + return result; + } + + // A list as a .dcf file; hashes sorted, so the same words always give the same bytes + inline std::string Serialize(const ChatFilterWords::WordList& list) { + std::vector hashes(list.hashes.begin(), list.hashes.end()); + std::sort(hashes.begin(), hashes.end()); + std::string out; + out.reserve(headerSize + hashes.size() * 8); + detail::WriteLE(out, header, 4); + detail::WriteLE(out, formatVersion, 4); + detail::WriteLE(out, list.maxWords, 4); + detail::WriteLE(out, hashes.size(), 8); + for (const auto hash : hashes) detail::WriteLE(out, hash, 8); + return out; + } + + // A plain block list: one word or phrase per line, normalized (ChatFilterWords::NormalizeEntry); empty lines skipped + inline ChatFilterWords::WordList BlockListFromText(std::string_view text) { + ChatFilterWords::WordList list; + size_t start = 0; + while (start < text.size()) { + auto end = text.find('\n', start); + if (end == std::string_view::npos) end = text.size(); + list.AddEntry(ChatFilterWords::NormalizeEntry(text.substr(start, end - start))); + start = end + 1; + } + return list; + } + + // A plain allow list (chatplus_en_us.txt): one word per line, lower case, compared whole (as the filter always has) + inline ChatFilterWords::WordList AllowListFromText(std::string_view text) { + ChatFilterWords::WordList list; + size_t start = 0; + while (start < text.size()) { + auto end = text.find('\n', start); + if (end == std::string_view::npos) end = text.size(); + std::string line(text.substr(start, end - start)); + std::erase(line, '\r'); + line = ChatFilterWords::AsciiLower(std::move(line)); + list.hashes.insert(ChatFilterWords::Hash(line)); + list.maxWords = std::max(list.maxWords, 1u); + start = end + 1; + } + return list; + } + + inline std::optional ReadBytes(const std::filesystem::path& path) { + std::ifstream in(path, std::ios::binary); + if (!in) return std::nullopt; + return std::string((std::istreambuf_iterator(in)), std::istreambuf_iterator()); + } + + inline ParseResult ReadFile(const std::filesystem::path& path) { + const auto bytes = ReadBytes(path); + if (!bytes) return { eStatus::MISSING }; + return Parse(*bytes); + } + + // Writes the file whole or not at all (a temporary file renamed over it), so servers starting together don't read half a file + inline bool WriteFile(const std::filesystem::path& path, const ChatFilterWords::WordList& list) { + auto temp = path; + temp += "." + std::to_string(std::random_device{}()) + ".tmp"; + { + std::ofstream out(temp, std::ios::binary | std::ios::trunc); + if (!out) return false; + const auto bytes = Serialize(list); + out.write(bytes.data(), static_cast(bytes.size())); + if (!out) return false; + } + std::error_code error; + std::filesystem::rename(temp, path, error); + if (!error) return true; + std::filesystem::remove(temp, error); + return false; + } +} diff --git a/dChatFilter/dChatFilter.cpp b/dChatFilter/dChatFilter.cpp index 83b837ece..fcbb4e92b 100644 --- a/dChatFilter/dChatFilter.cpp +++ b/dChatFilter/dChatFilter.cpp @@ -1,15 +1,8 @@ #include "dChatFilter.h" -#include "BinaryIO.h" -#include -#include -#include -#include -#include -#include -#include "dCommonVars.h" +#include + #include "Logger.h" -#include "dConfig.h" #include "Database.h" #include "Game.h" #include "eGameMasterLevel.h" @@ -19,154 +12,96 @@ using namespace dChatFilterDCF; dChatFilter::dChatFilter(const std::string& filepath, bool dontGenerateDCF) { m_DontGenerateDCF = dontGenerateDCF; - if (!BinaryIO::DoesFileExist(filepath + ".dcf") || m_DontGenerateDCF) { - ReadWordlistPlaintext(filepath + ".txt", true); - if (!m_DontGenerateDCF) ExportWordlistToDCF(filepath + ".dcf", true); - } else if (!ReadWordlistDCF(filepath + ".dcf", true)) { - ReadWordlistPlaintext(filepath + ".txt", true); - ExportWordlistToDCF(filepath + ".dcf", true); - } + LoadAllowList(filepath); + LoadBlockList(); - if (BinaryIO::DoesFileExist("blocklist.dcf")) { - ReadWordlistDCF("blocklist.dcf", false); - } - - //Read player names that are ok as well: - auto approvedNames = Database::Get()->GetApprovedCharacterNames(); - for (auto& name : approvedNames) { - std::transform(name.begin(), name.end(), name.begin(), ::tolower); //Transform to lowercase - m_ApprovedWords.push_back(CalculateHash(name)); + // Approved character names count as allowed words + for (const auto& name : Database::Get()->GetApprovedCharacterNames()) { + m_Lists.approved.hashes.insert(ChatFilterWords::Hash(ChatFilterWords::AsciiLower(name))); } ReloadCustomWords(); } -void dChatFilter::ReloadCustomWords() { - m_CustomAllowedWords.clear(); - m_CustomBlockedWords.clear(); - // Words remembered as not allowed may be allowed now - m_UserUnapprovedWordCache.clear(); - for (const auto& word : Database::Get()->GetChatFilterWords()) { - (word.allowed ? m_CustomAllowedWords : m_CustomBlockedWords).insert(CalculateHash(NormalizeWord(word.word))); - } -} - -dChatFilter::~dChatFilter() { - m_ApprovedWords.clear(); - m_DeniedWords.clear(); -} - -void dChatFilter::ReadWordlistPlaintext(const std::string& filepath, bool allowList) { - std::ifstream file(filepath); - if (file) { - std::string line; - while (std::getline(file, line)) { - line.erase(std::remove(line.begin(), line.end(), '\r'), line.end()); - std::transform(line.begin(), line.end(), line.begin(), ::tolower); //Transform to lowercase - if (allowList) m_ApprovedWords.push_back(CalculateHash(line)); - else m_DeniedWords.push_back(CalculateHash(line)); +void dChatFilter::LoadAllowList(const std::string& filepath) { + const std::string dcf = filepath + ".dcf"; + const std::string txt = filepath + ".txt"; + if (!m_DontGenerateDCF) { + auto cached = ReadFile(dcf); + if (cached.status == eStatus::OK) { + m_Lists.approved = std::move(cached.list); + return; } + if (cached.status != eStatus::MISSING) LOG("%s is %s; building it again from %s", dcf.c_str(), StatusText(cached.status), txt.c_str()); } + + const auto text = ReadBytes(txt); + if (!text) { + LOG("Could not read the chat filter's allowed words (%s)", txt.c_str()); + return; + } + m_Lists.approved = AllowListFromText(*text); + if (!m_DontGenerateDCF && !WriteFile(dcf, m_Lists.approved)) LOG("Could not write %s", dcf.c_str()); } -bool dChatFilter::ReadWordlistDCF(const std::string& filepath, bool allowList) { - std::ifstream file(filepath, std::ios::binary); - if (file) { - fileHeader hdr; - BinaryIO::BinaryRead(file, hdr); - if (hdr.header != header) { - file.close(); - return false; +void dChatFilter::LoadBlockList() { + std::error_code error; + const bool hasText = std::filesystem::exists(BLOCK_LIST_TEXT, error); + if (hasText) { + const auto text = ReadBytes(BLOCK_LIST_TEXT); + const auto existing = ReadFile(BLOCK_LIST_FILE); + // Rebuilt when the .dcf is missing, unreadable or older than the text + bool stale = existing.status != eStatus::OK; + if (!stale) { + std::error_code textError, fileError; + const auto textTime = std::filesystem::last_write_time(BLOCK_LIST_TEXT, textError); + const auto fileTime = std::filesystem::last_write_time(BLOCK_LIST_FILE, fileError); + stale = textError || fileError || fileTime < textTime; } - - if (hdr.formatVersion == formatVersion) { - size_t wordsToRead = 0; - BinaryIO::BinaryRead(file, wordsToRead); - if (allowList) m_ApprovedWords.reserve(wordsToRead); - else m_DeniedWords.reserve(wordsToRead); - - size_t word = 0; - for (size_t i = 0; i < wordsToRead; ++i) { - BinaryIO::BinaryRead(file, word); - if (allowList) m_ApprovedWords.push_back(word); - else m_DeniedWords.push_back(word); + if (text && (m_DontGenerateDCF || stale)) { + m_Lists.denied = BlockListFromText(*text); + if (m_DontGenerateDCF) { + LOG("Loaded %zu blocked words and phrases from %s", m_Lists.denied.Size(), BLOCK_LIST_TEXT); + return; } - - return true; - } else { - file.close(); - return false; + if (WriteFile(BLOCK_LIST_FILE, m_Lists.denied)) { + LOG("Built %s from %s (%zu words and phrases)", BLOCK_LIST_FILE, BLOCK_LIST_TEXT, m_Lists.denied.Size()); + } else { + LOG("Could not write %s", BLOCK_LIST_FILE); + } + return; } } - return false; + auto blocked = ReadFile(BLOCK_LIST_FILE); + switch (blocked.status) { + case eStatus::OK: + m_Lists.denied = std::move(blocked.list); + break; + case eStatus::MISSING: + LOG("No %s: best friends' free chat stops every message. Put the blocked words in %s next to the servers (one word or phrase per line) and start the servers again.", + BLOCK_LIST_FILE, BLOCK_LIST_TEXT); + break; + case eStatus::OLD_FORMAT: + LOG("%s is in the old format (version 2), whose hashes depend on the compiler and platform, so it can't be read; best friends' free chat stops every message. " + "Put the plain word list in %s next to the servers (one word or phrase per line) and start the servers again to rebuild it.", + BLOCK_LIST_FILE, BLOCK_LIST_TEXT); + break; + default: + LOG("%s is %s and can't be read; best friends' free chat stops every message. Rebuild it from %s.", BLOCK_LIST_FILE, StatusText(blocked.status), BLOCK_LIST_TEXT); + break; + } } -void dChatFilter::ExportWordlistToDCF(const std::string& filepath, bool allowList) { - std::ofstream file(filepath, std::ios::binary | std::ios_base::out); - if (file) { - BinaryIO::BinaryWrite(file, uint32_t(dChatFilterDCF::header)); - BinaryIO::BinaryWrite(file, uint32_t(dChatFilterDCF::formatVersion)); - BinaryIO::BinaryWrite(file, size_t(allowList ? m_ApprovedWords.size() : m_DeniedWords.size())); - - for (size_t word : allowList ? m_ApprovedWords : m_DeniedWords) { - BinaryIO::BinaryWrite(file, word); - } - - file.close(); +void dChatFilter::ReloadCustomWords() { + m_Lists.customAllowed = {}; + m_Lists.customBlocked = {}; + for (const auto& word : Database::Get()->GetChatFilterWords()) { + (word.allowed ? m_Lists.customAllowed : m_Lists.customBlocked).AddEntry(ChatFilterWords::NormalizeEntry(word.word)); } } std::set> dChatFilter::IsSentenceOkay(const std::string& message, eGameMasterLevel gmLevel, bool allowList) { if (gmLevel > eGameMasterLevel::FORUM_MODERATOR) return { }; //If anything but a forum mod, return true. - if (message.empty()) return { }; - if (!allowList && m_DeniedWords.empty()) return { { 0, message.length() } }; - - std::stringstream sMessage(message); - std::string segment; - - std::set> listOfBadSegments; - - uint32_t position = 0; - - while (std::getline(sMessage, segment, ' ')) { - std::string originalSegment = segment; - - segment = NormalizeWord(segment); - - size_t hash = CalculateHash(segment); - - // Blocked on the dashboard: stopped in every kind of chat - if (m_CustomBlockedWords.contains(hash)) { - listOfBadSegments.emplace(position, originalSegment.length()); - position += originalSegment.length() + 1; - continue; - } - - if (std::find(m_UserUnapprovedWordCache.begin(), m_UserUnapprovedWordCache.end(), hash) != m_UserUnapprovedWordCache.end() && allowList) { - listOfBadSegments.emplace(position, originalSegment.length()); - } - - if (std::find(m_ApprovedWords.begin(), m_ApprovedWords.end(), hash) == m_ApprovedWords.end() && !m_CustomAllowedWords.contains(hash) && allowList) { - m_UserUnapprovedWordCache.push_back(hash); - listOfBadSegments.emplace(position, originalSegment.length()); - } - - if (std::find(m_DeniedWords.begin(), m_DeniedWords.end(), hash) != m_DeniedWords.end() && !allowList) { - m_UserUnapprovedWordCache.push_back(hash); - listOfBadSegments.emplace(position, originalSegment.length()); - } - - position += originalSegment.length() + 1; - } - - return listOfBadSegments; -} - -size_t dChatFilter::CalculateHash(const std::string& word) { - std::hash hash{}; - - size_t value = hash(word); - - return value; + return ChatFilterWords::CheckMessage(message, allowList, m_Lists); } diff --git a/dChatFilter/dChatFilter.h b/dChatFilter/dChatFilter.h index cd8dac353..e6a221df5 100644 --- a/dChatFilter/dChatFilter.h +++ b/dChatFilter/dChatFilter.h @@ -1,70 +1,52 @@ #pragma once #include #include +#include +#include #include -#include -#include #include +#include +#include "ChatFilterCore.h" #include "dCommonVars.h" enum class eGameMasterLevel : uint8_t; -namespace dChatFilterDCF { - static const uint32_t header = ('D' + ('C' << 8) + ('F' << 16) + ('B' << 24)); - static const uint32_t formatVersion = 2; - - struct fileHeader { - uint32_t header; - uint32_t formatVersion; - }; -}; class dChatFilter { public: + /** + * Loads the allow list (filepath + ".txt", cached as filepath + ".dcf") and the block list (blocklist.dcf next to the + * servers, rebuilt from blocklist.txt there when that file is newer). dontGenerateDCF: read the plain lists only and + * write no .dcf files. + */ dChatFilter(const std::string& filepath, bool dontGenerateDCF); - ~dChatFilter(); + ~dChatFilter() = default; - void ReadWordlistPlaintext(const std::string& filepath, bool allowList); - bool ReadWordlistDCF(const std::string& filepath, bool allowList); - void ExportWordlistToDCF(const std::string& filepath, bool allowList); std::set> IsSentenceOkay(const std::string& message, eGameMasterLevel gmLevel, bool allowList = true); // Whether a deny list is loaded (without one, IsSentenceOkay(..., false) refuses every message) - bool HasDenyList() const { return !m_DeniedWords.empty(); } + bool HasDenyList() const { return !m_Lists.denied.Empty(); } /** * Load the words staff added on the dashboard (chat_filter_words) again, replacing the ones loaded before. - * Allowed words are accepted in whitelisted chat; blocked words are always stopped, even when a file allows them. + * Allowed words are accepted in whitelisted chat; blocked words and phrases are always stopped, even when a file allows them. */ void ReloadCustomWords(); // A word as the filter compares it: lower case, without ! ? ; . , - static std::string NormalizeWord(std::string word) { - std::erase_if(word, [](char c) { return c == '!' || c == '?' || c == ';' || c == '.' || c == ','; }); - std::transform(word.begin(), word.end(), word.begin(), ::tolower); //Transform to lowercase - return word; - } + static std::string NormalizeWord(std::string word) { return ChatFilterWords::NormalizeWord(std::move(word)); } // A message split into words (at spaces) the way the filter checks it static std::vector Words(const std::string& message) { std::vector words; - size_t start = 0; - while (start <= message.size()) { - const auto end = std::min(message.find(' ', start), message.size()); - words.push_back(NormalizeWord(message.substr(start, end - start))); - start = end + 1; - } + for (auto& token : ChatFilterWords::Tokenize(message)) words.push_back(std::move(token.word)); return words; } private: - bool m_DontGenerateDCF; - std::vector m_DeniedWords; - std::vector m_ApprovedWords; - std::vector m_UserUnapprovedWordCache; - std::unordered_set m_CustomAllowedWords; - std::unordered_set m_CustomBlockedWords; + void LoadAllowList(const std::string& filepath); + void LoadBlockList(); - //Private functions: - size_t CalculateHash(const std::string& word); + bool m_DontGenerateDCF; + ChatFilterWords::Lists m_Lists; }; diff --git a/dDashboardServer/routes/ChatFilterWords.h b/dDashboardServer/routes/ChatFilterWords.h index 29b07cea8..00dcb4cb4 100644 --- a/dDashboardServer/routes/ChatFilterWords.h +++ b/dDashboardServer/routes/ChatFilterWords.h @@ -8,23 +8,26 @@ #include #include -#include "dChatFilter.h" +#include "ChatFilterCore.h" // Words for the chat filter page, compared the way the chat filter compares them (see ModerationTools.h) namespace ModerationTools { constexpr size_t MAX_FILTER_WORD = 64; - // A word staff typed for the chat filter, as the filter compares it (lower case, no ! ? ; . ,); nullopt if it isn't - // one word of 1-64 characters. Pure; unit tested. + // A word or phrase staff typed for the chat filter, as the filter compares it (lower case, no ! ? ; . ,, words joined by + // one space); nullopt if it isn't 1-64 characters or has no word. Pure; unit tested. inline std::optional FilterWord(std::string text) { text.erase(0, text.find_first_not_of(" \t\r\n")); text.erase(text.find_last_not_of(" \t\r\n") + 1); - if (text.empty() || text.size() > MAX_FILTER_WORD || text.find_first_of(" \t\r\n") != std::string::npos) return std::nullopt; - auto word = dChatFilter::NormalizeWord(text); - if (word.empty()) return std::nullopt; - return word; + if (text.empty() || text.size() > MAX_FILTER_WORD) return std::nullopt; + auto entry = ChatFilterWords::NormalizeEntry(text); + if (entry.empty()) return std::nullopt; + return entry; } + // Whether an entry is a phrase (more than one word). Phrases can only be blocked: whitelist chat checks one word at a time. + inline bool IsPhrase(const std::string& entry) { return ChatFilterWords::WordCount(entry) > 1; } + // The words of a plain word list (chatplus_en_us.txt) as the filter reads them: one per line, lower case; sorted, each once inline std::vector FileWords(const std::string& text) { std::vector words; @@ -33,7 +36,7 @@ namespace ModerationTools { const auto end = std::min(text.find('\n', start), text.size()); auto line = text.substr(start, end - start); std::erase(line, '\r'); - std::transform(line.begin(), line.end(), line.begin(), ::tolower); + line = ChatFilterWords::AsciiLower(std::move(line)); if (!line.empty()) words.push_back(std::move(line)); start = end + 1; } @@ -42,28 +45,12 @@ namespace ModerationTools { return words; } - // The hashes of a .dcf word list (blocklist.dcf), as dChatFilter::ReadWordlistDCF reads them; nullopt if it isn't one - inline std::optional> DcfHashes(const std::string& bytes) { - dChatFilterDCF::fileHeader header{}; - size_t count = 0; - if (bytes.size() < sizeof(header) + sizeof(count)) return std::nullopt; - std::memcpy(&header, bytes.data(), sizeof(header)); - if (header.header != dChatFilterDCF::header || header.formatVersion != dChatFilterDCF::formatVersion) return std::nullopt; - std::memcpy(&count, bytes.data() + sizeof(header), sizeof(count)); - const size_t offset = sizeof(header) + sizeof(count); - if (count > (bytes.size() - offset) / sizeof(size_t)) return std::nullopt; - std::vector hashes(count); - if (count) std::memcpy(hashes.data(), bytes.data() + offset, count * sizeof(size_t)); - return hashes; - } - - // A word's hash as the filter stores it (dChatFilter::CalculateHash) - inline size_t WordHash(const std::string& word) { return std::hash{}(word); } - - // Whether a message contains `word` as one of the words the chat filter checks - inline bool HasFilterWord(const std::string& message, const std::string& word) { - const auto words = dChatFilter::Words(message); - return std::find(words.begin(), words.end(), word) != words.end(); + // Whether a message contains a word or phrase (FilterWord) as the chat filter reads it: whole words, in a row, skipping + // pieces that are only punctuation + inline bool HasFilterWord(const std::string& message, const std::string& entry) { + const auto tokens = ChatFilterWords::Tokenize(message); + const auto words = ChatFilterWords::WordCount(entry); + return !ChatFilterWords::FindBlocked(tokens, words, [&entry](const std::string& run) { return run == entry; }).empty(); } // What the filter decides about one word of a message, and why @@ -72,6 +59,7 @@ namespace ModerationTools { std::string word; // as the filter compares it bool stopped{}; std::string reason; // blocked_here, allowed_here, allow_file, character_name, not_allowed, block_file, not_in_block_file, no_block_file + std::string phrase; // the blocked phrase this word is part of (blocked_here or block_file), when it was a phrase }; // Where the filter finds its words (callbacks keep this pure; the route reads the files and the database) @@ -81,33 +69,49 @@ namespace ModerationTools { std::function characterName; // approved character names count as allowed words std::function blockFile; // blocklist.dcf (by hash) bool blockFileLoaded{}; + uint32_t maxWords{ 1 }; // the longest blocked phrase, in words (here or in the file) }; /** * Each word of a message with what dChatFilter::IsSentenceOkay decides about it for a player below GM level 2 (higher - * levels skip the filter). Normal chat (allowList) needs every word allowed; best friends' free chat (!allowList) stops - * only blocked words, or every word when there is no blocked words file. Words are split at spaces as the filter - * splits them. Pure; unit tested. + * levels skip the filter). Blocked words and phrases (here always, the block file's in free chat) are stopped; a phrase + * stops each of its words. Normal chat (allowList) needs every other word allowed, one at a time; best friends' free + * chat stops only blocked ones, or every word when there is no blocked words file. Words are split at spaces as the + * filter splits them (ChatFilterWords::CheckMessage). Pure; unit tested. */ inline std::vector ExplainMessage(const std::string& message, bool allowList, const WordSources& sources) { + const auto tokens = ChatFilterWords::Tokenize(message); std::vector verdicts; - std::stringstream stream(message); - std::string segment; - while (std::getline(stream, segment, ' ')) { - WordVerdict verdict{ segment, dChatFilter::NormalizeWord(segment) }; - const auto here = sources.dashboard(verdict.word); - if (!allowList && !sources.blockFileLoaded) { + for (const auto& token : tokens) verdicts.push_back({ message.substr(token.position, token.length), token.word }); + if (!allowList && !sources.blockFileLoaded) { + for (auto& verdict : verdicts) { verdict.stopped = true; verdict.reason = "no_block_file"; - } else if (here && !*here) { - verdict.stopped = true; - verdict.reason = "blocked_here"; - } else if (!allowList) { - verdict.stopped = sources.blockFile(verdict.word); - verdict.reason = verdict.stopped ? "block_file" : "not_in_block_file"; + } + return verdicts; + } + + const auto blockedHere = [&sources](const std::string& entry) { const auto here = sources.dashboard(entry); return here && !*here; }; + const auto matches = ChatFilterWords::FindBlocked(tokens, std::max(sources.maxWords, 1u), [&](const std::string& entry) { + return blockedHere(entry) || (!allowList && sources.blockFile(entry)); + }); + for (const auto& match : matches) { + const auto reason = blockedHere(match.entry) ? "blocked_here" : "block_file"; + for (size_t i = match.first; i <= match.last; i++) { + verdicts[i].stopped = true; + verdicts[i].reason = reason; + if (ChatFilterWords::WordCount(match.entry) > 1) verdicts[i].phrase = match.entry; + } + } + + for (auto& verdict : verdicts) { + if (verdict.stopped) continue; + const auto here = sources.dashboard(verdict.word); + if (!allowList) { + verdict.reason = "not_in_block_file"; } else if (sources.allowFile(verdict.word)) { verdict.reason = "allow_file"; - } else if (here) { + } else if (here && *here) { verdict.reason = "allowed_here"; } else if (sources.characterName(verdict.word)) { verdict.reason = "character_name"; @@ -115,7 +119,6 @@ namespace ModerationTools { verdict.stopped = true; verdict.reason = "not_allowed"; } - verdicts.push_back(std::move(verdict)); } return verdicts; } diff --git a/dDashboardServer/routes/ModerationTools.cpp b/dDashboardServer/routes/ModerationTools.cpp index 517288fad..31c63e9de 100644 --- a/dDashboardServer/routes/ModerationTools.cpp +++ b/dDashboardServer/routes/ModerationTools.cpp @@ -31,7 +31,8 @@ namespace { // The chat filter's files, as the servers load them: the allowed words from the client's res folder, the blocked // words (only their hashes) next to the servers constexpr const char* ALLOW_FILE = "chatplus_en_us.txt"; - constexpr const char* BLOCK_FILE = "blocklist.dcf"; + constexpr const char* BLOCK_FILE = dChatFilterDCF::BLOCK_LIST_FILE; + constexpr const char* BLOCK_TEXT = dChatFilterDCF::BLOCK_LIST_TEXT; constexpr uint32_t FILE_WORDS_PAGE = 200; // Recent chat searched when checking what a word would change constexpr uint32_t CHECK_MESSAGES = 1000; @@ -216,11 +217,8 @@ namespace { return text ? ModerationTools::FileWords(*text) : std::vector{}; } - std::optional> BlockFileHashes() { - std::ifstream in(BLOCK_FILE, std::ios::binary); - if (!in) return std::nullopt; - return ModerationTools::DcfHashes(std::string((std::istreambuf_iterator(in)), std::istreambuf_iterator())); - } + // blocklist.dcf as the servers read it (an old-format file is refused, as they refuse it) + dChatFilterDCF::ParseResult BlockFile() { return dChatFilterDCF::ReadFile(BLOCK_FILE); } // The dashboard's own lists by word: true allowed, false blocked std::map DashboardWords() { @@ -258,25 +256,27 @@ namespace { const auto it = dashboard.find(word); words.push_back({ {"word", word}, {"dashboard", it == dashboard.end() ? nlohmann::json(nullptr) : nlohmann::json(it->second ? "allowed" : "blocked")} }); } - const auto blocked = BlockFileHashes(); + const auto blocked = BlockFile(); uint32_t imported = 0; for (const auto& word : all) if (dashboard.contains(word)) imported++; JsonSuccess(reply, { {"allowFile", ALLOW_FILE}, {"allowFileFound", !all.empty()}, {"allowTotal", all.size()}, {"matched", matched}, {"start", start}, {"pageSize", FILE_WORDS_PAGE}, {"words", words}, {"onDashboard", imported}, - {"blockFile", BLOCK_FILE}, {"blockFileFound", blocked.has_value()}, {"blockTotal", blocked ? blocked->size() : 0} }); + {"blockFile", BLOCK_FILE}, {"blockFileFound", blocked.status == dChatFilterDCF::eStatus::OK}, {"blockTotal", blocked.list.Size()}, + {"blockMaxWords", blocked.list.maxWords}, {"blockFileStatus", dChatFilterDCF::StatusText(blocked.status)}, + {"blockFileOld", blocked.status == dChatFilterDCF::eStatus::OLD_FORMAT}, {"blockText", BLOCK_TEXT} }); }); Route(eHTTPMethod::GET, "/api/chat_filter/lookup", Perm("chat_filter_manage"), - "Where a word stands: in the allowed words file, in the blocked words file (by its hash), and on the dashboard's lists. Query: word", + "Where a word or phrase stands: in the allowed words file, in the blocked words file (by its hash), and on the dashboard's lists. Query: word", [](HTTPReply& reply, const HTTPContext& context) { const auto word = ModerationTools::FilterWord(QueryValue(context.queryString, "word")); - if (!word) return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Type one word (no spaces), up to 64 characters"); + if (!word) return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Type a word or phrase, up to 64 characters"); const auto all = AllowFileWords(); - const auto blocked = BlockFileHashes(); + const auto blocked = BlockFile(); const auto dashboard = DashboardWords(); const auto it = dashboard.find(*word); - JsonSuccess(reply, { {"word", *word}, {"inAllowFile", std::binary_search(all.begin(), all.end(), *word)}, - {"inBlockFile", blocked && std::find(blocked->begin(), blocked->end(), ModerationTools::WordHash(*word)) != blocked->end()}, + JsonSuccess(reply, { {"word", *word}, {"phrase", ModerationTools::IsPhrase(*word)}, {"inAllowFile", std::binary_search(all.begin(), all.end(), *word)}, + {"inBlockFile", blocked.list.Contains(*word)}, {"dashboard", it == dashboard.end() ? nlohmann::json(nullptr) : nlohmann::json(it->second ? "allowed" : "blocked")} }); }); @@ -293,7 +293,7 @@ namespace { for (const auto& word : all) { // Only words the filter could compare (the file's odd lines with spaces or punctuation stay in the file only) const auto filterWord = ModerationTools::FilterWord(word); - if (!filterWord || *filterWord != word || dashboard.contains(word)) continue; + if (!filterWord || *filterWord != word || ModerationTools::IsPhrase(word) || dashboard.contains(word)) continue; Database::Get()->SetChatFilterWord({ word, true, context.authenticatedUser, now }); added++; } @@ -305,13 +305,17 @@ namespace { }); Route(eHTTPMethod::POST, "/api/chat_filter/words", Perm("chat_filter_manage"), - "Allow or block a word (or move it to the other list); running worlds pick it up at once. Body: {word, allowed: bool}", + "Allow or block a word, or block a phrase (or move it to the other list); running worlds pick it up at once. Phrases can't be allowed: " + "whitelist chat checks each word on its own, as the client does. Body: {word, allowed: bool}", [](HTTPReply& reply, const HTTPContext& context) { const auto body = ParseBody(context); if (!body) return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Invalid JSON"); const auto word = ModerationTools::FilterWord(body->value("word", "")); - if (!word) return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Type one word (no spaces), up to 64 characters"); + if (!word) return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Type a word or phrase, up to 64 characters"); const bool allowed = body->value("allowed", false); + if (allowed && ModerationTools::IsPhrase(*word)) { + return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Phrases can only be blocked: normal chat checks each word on its own, so allow the words instead"); + } Database::Get()->SetChatFilterWord({ *word, allowed, context.authenticatedUser, static_cast(std::time(nullptr)) }); Audit(context, allowed ? "chat_filter_allow" : "chat_filter_block", (allowed ? "Allowed \"" : "Blocked \"") + *word + "\" in chat"); BroadcastTableChanged("chat_filter"); @@ -345,12 +349,11 @@ namespace { if (message.empty() || message.size() > 300) return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Type a message of 1 to 300 characters"); const bool allowList = QueryValue(context.queryString, "chat") != "free"; const auto all = AllowFileWords(); - const auto blocked = BlockFileHashes(); + const auto blocked = BlockFile(); const auto dashboard = DashboardWords(); std::set names; for (auto name : Database::Get()->GetApprovedCharacterNames()) { - std::transform(name.begin(), name.end(), name.begin(), ::tolower); - names.insert(std::move(name)); + names.insert(ChatFilterWords::AsciiLower(std::move(name))); } ModerationTools::WordSources sources; sources.dashboard = [&dashboard](const std::string& word) -> std::optional { @@ -359,18 +362,18 @@ namespace { }; sources.allowFile = [&all](const std::string& word) { return std::binary_search(all.begin(), all.end(), word); }; sources.characterName = [&names](const std::string& word) { return names.contains(word); }; - sources.blockFile = [&blocked](const std::string& word) { - return blocked && std::find(blocked->begin(), blocked->end(), ModerationTools::WordHash(word)) != blocked->end(); - }; - sources.blockFileLoaded = blocked && !blocked->empty(); + sources.blockFile = [&blocked](const std::string& entry) { return blocked.list.Contains(entry); }; + sources.blockFileLoaded = !blocked.list.Empty(); + sources.maxWords = blocked.list.maxWords; + for (const auto& [entry, allowed] : dashboard) if (!allowed) sources.maxWords = std::max(sources.maxWords, ChatFilterWords::WordCount(entry)); nlohmann::json words = nlohmann::json::array(); bool stopped = false; for (const auto& verdict : ModerationTools::ExplainMessage(message, allowList, sources)) { stopped |= verdict.stopped; - words.push_back({ {"text", verdict.text}, {"word", verdict.word}, {"stopped", verdict.stopped}, {"reason", verdict.reason} }); + words.push_back({ {"text", verdict.text}, {"word", verdict.word}, {"stopped", verdict.stopped}, {"reason", verdict.reason}, {"phrase", verdict.phrase} }); } JsonSuccess(reply, { {"message", message}, {"chat", allowList ? "normal" : "free"}, {"stopped", stopped}, {"words", words}, - {"allowFileFound", !all.empty()}, {"blockFileFound", blocked.has_value()} }); + {"allowFileFound", !all.empty()}, {"blockFileFound", blocked.status == dChatFilterDCF::eStatus::OK} }); }); Route(eHTTPMethod::GET, "/api/chat_filter/check", Perm("chat_filter_manage"), @@ -379,7 +382,7 @@ namespace { [](HTTPReply& reply, const HTTPContext& context) { if (!Can(context, "chat_view")) return JsonError(reply, eHTTPStatusCode::FORBIDDEN, "Reading chat needs the chat_view permission"); const auto word = ModerationTools::FilterWord(QueryValue(context.queryString, "word")); - if (!word) return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Type one word (no spaces), up to 64 characters"); + if (!word) return JsonError(reply, eHTTPStatusCode::BAD_REQUEST, "Type a word or phrase, up to 64 characters"); const bool allowed = QueryValue(context.queryString, "allowed") == "1"; IChatLog::ChatQuery query; query.search = *word; diff --git a/tests/dCommonTests/CMakeLists.txt b/tests/dCommonTests/CMakeLists.txt index bca761760..772126890 100644 --- a/tests/dCommonTests/CMakeLists.txt +++ b/tests/dCommonTests/CMakeLists.txt @@ -39,6 +39,7 @@ set(DCOMMONTEST_SOURCES "FdbReaderTests.cpp" "FdbSnapshotTests.cpp" "WorldFileWatchTests.cpp" + "ChatFilterCoreTests.cpp" ) add_subdirectory(dEnumsTests) @@ -59,6 +60,8 @@ endif() target_link_libraries(dCommonTests ${COMMON_LIBRARIES} MD5 GTest::gtest_main) # SpareBackoff.h (header only) target_include_directories(dCommonTests PRIVATE "${PROJECT_SOURCE_DIR}/dMasterServer") +# ChatFilterCore.h (header only) +target_include_directories(dCommonTests PRIVATE "${PROJECT_SOURCE_DIR}/dChatFilter") # Copy test files to testing directory add_subdirectory(TestBitStreams) diff --git a/tests/dCommonTests/ChatFilterCoreTests.cpp b/tests/dCommonTests/ChatFilterCoreTests.cpp new file mode 100644 index 000000000..50b66590b --- /dev/null +++ b/tests/dCommonTests/ChatFilterCoreTests.cpp @@ -0,0 +1,149 @@ +#include + +#include +#include + +#include "ChatFilterCore.h" + +using namespace ChatFilterWords; +using Spans = std::set>; + +// The hash is a compile-time constant: the same value on every compiler, standard library and platform +static_assert(Hash("") == 0xcbf29ce484222325ULL); +static_assert(Hash("a") == 0xaf63dc4c8601ec8cULL); +static_assert(Hash("foobar") == 0x85944171f73967e8ULL); + +TEST(ChatFilterCoreTest, HashIsFnv1a64) { + // The standard FNV-1a 64 test vectors + EXPECT_EQ(Hash(""), 0xcbf29ce484222325ULL); + EXPECT_EQ(Hash("a"), 0xaf63dc4c8601ec8cULL); + EXPECT_EQ(Hash("foobar"), 0x85944171f73967e8ULL); + // Words of the client's chatplus_en_us.txt, as the filter hashes them (lower case) + EXPECT_EQ(Hash("hello"), 0xa430d84680aabd0bULL); + EXPECT_EQ(Hash("brick"), 0xf9236d1e24832c9aULL); + EXPECT_EQ(Hash(AsciiLower("Brick")), Hash("brick")); + // A phrase is its words joined by one space + EXPECT_EQ(Hash("bad phrase"), 0x72e9fce49c1b0875ULL); + // Bytes, not chars: a high byte hashes as 0x80-0xFF whether char is signed or not + EXPECT_EQ(Hash("\xC3\xA9"), 0x0ac21707b7181e01ULL); +} + +TEST(ChatFilterCoreTest, NormalizeEntry) { + EXPECT_EQ(NormalizeWord("Hello!?"), "hello"); + EXPECT_EQ(NormalizeEntry(" Bad \t PHRASE! "), "bad phrase"); + EXPECT_EQ(NormalizeEntry("word , here"), "word here"); + EXPECT_EQ(NormalizeEntry("..."), ""); + EXPECT_EQ(WordCount("bad phrase here"), 3u); + EXPECT_EQ(WordCount("word"), 1u); + EXPECT_EQ(WordCount(""), 0u); + // ASCII only: other bytes are left alone whatever the locale + EXPECT_EQ(AsciiLower("\xC3\x89Z"), "\xC3\x89z"); +} + +TEST(ChatFilterCoreTest, DcfBytesAreFixed) { + WordList list; + list.AddEntry("a"); + list.AddEntry("bad phrase"); + const auto bytes = dChatFilterDCF::Serialize(list); + // magic DCFB, version 3, 2 words at most, 2 hashes, then the hashes sorted, all little-endian + const std::string expected( + "DCFB" "\x03\x00\x00\x00" "\x02\x00\x00\x00" "\x02\x00\x00\x00\x00\x00\x00\x00" + "\x75\x08\x1b\x9c\xe4\xfc\xe9\x72" "\x8c\xec\x01\x86\x4c\xdc\x63\xaf", 4 + 4 + 4 + 8 + 16); + EXPECT_EQ(bytes, expected); + + const auto parsed = dChatFilterDCF::Parse(bytes); + ASSERT_EQ(parsed.status, dChatFilterDCF::eStatus::OK); + EXPECT_EQ(parsed.list.maxWords, 2u); + EXPECT_EQ(parsed.list.hashes, list.hashes); +} + +TEST(ChatFilterCoreTest, DcfRejectsOldAndBadFiles) { + // Version 2 (std::hash, platform dependent) is refused, not guessed at + const std::string old("DCFB" "\x02\x00\x00\x00" "\x01\x00\x00\x00\x00\x00\x00\x00" "\x11\x22\x33\x44\x55\x66\x77\x88", 24); + EXPECT_EQ(dChatFilterDCF::Parse(old).status, dChatFilterDCF::eStatus::OLD_FORMAT); + EXPECT_EQ(dChatFilterDCF::Parse("XCFB\x03\x00\x00\x00").status, dChatFilterDCF::eStatus::NOT_DCF); + EXPECT_EQ(dChatFilterDCF::Parse(std::string("DCFB\x09\x00\x00\x00", 8)).status, dChatFilterDCF::eStatus::UNKNOWN); + WordList list; + list.AddEntry("word"); + auto bytes = dChatFilterDCF::Serialize(list); + bytes.pop_back(); + EXPECT_EQ(dChatFilterDCF::Parse(bytes).status, dChatFilterDCF::eStatus::TRUNCATED); + EXPECT_EQ(dChatFilterDCF::ReadFile("no such blocklist.dcf").status, dChatFilterDCF::eStatus::MISSING); +} + +TEST(ChatFilterCoreTest, BlockListFromText) { + const auto list = dChatFilterDCF::BlockListFromText("Badword\r\n\r\n Very BAD phrase!\nbadword\n...\n"); + EXPECT_EQ(list.Size(), 2u); + EXPECT_TRUE(list.Contains("badword")); + EXPECT_TRUE(list.Contains("very bad phrase")); + EXPECT_EQ(list.maxWords, 3u); + // Written and read back, the same list + const auto parsed = dChatFilterDCF::Parse(dChatFilterDCF::Serialize(list)); + ASSERT_EQ(parsed.status, dChatFilterDCF::eStatus::OK); + EXPECT_EQ(parsed.list.hashes, list.hashes); + EXPECT_EQ(parsed.list.maxWords, 3u); +} + +namespace { + Lists FreeChatLists(std::string_view blockText) { + Lists lists; + // Through the file format, as the servers load it + lists.denied = dChatFilterDCF::Parse(dChatFilterDCF::Serialize(dChatFilterDCF::BlockListFromText(blockText))).list; + lists.approved = dChatFilterDCF::AllowListFromText("hello\nthere\nfriend\nbad\nphrase\n"); + return lists; + } +} + +TEST(ChatFilterCoreTest, BlockedWordFromFileIsStopped) { + const auto lists = FreeChatLists("badword\n"); + EXPECT_EQ(CheckMessage("hello badword there", false, lists), (Spans{ { 6, 7 } })); + EXPECT_EQ(CheckMessage("hello BadWord!", false, lists), (Spans{ { 6, 8 } })); + EXPECT_TRUE(CheckMessage("hello there", false, lists).empty()); + // Whitelist chat doesn't use the block list: the word is simply not allowed there + EXPECT_EQ(CheckMessage("hello badword", true, lists), (Spans{ { 6, 7 } })); + // No block list: free chat stops everything + EXPECT_EQ(CheckMessage("hello", false, Lists{}), (Spans{ { 0, 5 } })); +} + +TEST(ChatFilterCoreTest, PhrasesAtStartMiddleEnd) { + const auto lists = FreeChatLists("bad phrase\n"); + EXPECT_EQ(CheckMessage("bad phrase hello", false, lists), (Spans{ { 0, 10 } })); + EXPECT_EQ(CheckMessage("hello bad phrase there", false, lists), (Spans{ { 6, 10 } })); + EXPECT_EQ(CheckMessage("hello bad phrase", false, lists), (Spans{ { 6, 10 } })); + // The words alone are fine + EXPECT_TRUE(CheckMessage("bad hello phrase", false, lists).empty()); + EXPECT_TRUE(CheckMessage("phrase bad", false, lists).empty()); +} + +TEST(ChatFilterCoreTest, PhrasesWithExtraSpacesAndPunctuation) { + const auto lists = FreeChatLists("bad phrase\n"); + // Two spaces: the span covers both words and the gap + EXPECT_EQ(CheckMessage("hi Bad Phrase!", false, lists), (Spans{ { 4, 12 } })); + // A piece that is only punctuation is skipped + EXPECT_EQ(CheckMessage("bad ... phrase", false, lists), (Spans{ { 0, 14 } })); + EXPECT_EQ(CheckMessage("bad, phrase.", false, lists), (Spans{ { 0, 12 } })); +} + +TEST(ChatFilterCoreTest, PhrasesOverlapWithWords) { + const auto lists = FreeChatLists("bad phrase\nphrase here\nbadword\nfriend\n"); + // Two phrases sharing a word become one span + EXPECT_EQ(CheckMessage("a bad phrase here b", false, lists), (Spans{ { 2, 15 } })); + // A blocked word inside a blocked phrase: one span for the phrase + EXPECT_EQ(CheckMessage("bad phrase", false, FreeChatLists("bad phrase\nphrase\n")), (Spans{ { 0, 10 } })); + // A blocked word right after a phrase: its own span + EXPECT_EQ(CheckMessage("bad phrase badword", false, lists), (Spans{ { 0, 10 }, { 11, 7 } })); + // Longest match wins where it starts + EXPECT_EQ(CheckMessage("friend bad phrase", false, FreeChatLists("friend\nfriend bad\nbad phrase\n")), (Spans{ { 0, 17 } })); +} + +TEST(ChatFilterCoreTest, DashboardPhrasesInWhitelistChat) { + auto lists = FreeChatLists(""); + lists.customBlocked.AddEntry(NormalizeEntry("Bad Phrase")); + // Blocked on the dashboard: stopped in whitelist chat too, even though each word is allowed + EXPECT_EQ(CheckMessage("hello bad phrase", true, lists), (Spans{ { 6, 10 } })); + EXPECT_TRUE(CheckMessage("hello bad there phrase", true, lists).empty()); + // Whitelist chat checks one word at a time, as the client does + EXPECT_EQ(CheckMessage("hello stranger", true, lists), (Spans{ { 6, 8 } })); + lists.customAllowed.AddEntry("stranger"); + EXPECT_TRUE(CheckMessage("hello stranger", true, lists).empty()); +} diff --git a/tests/dWebTests/ModerationToolsTests.cpp b/tests/dWebTests/ModerationToolsTests.cpp index 960f832c1..ccc2b980c 100644 --- a/tests/dWebTests/ModerationToolsTests.cpp +++ b/tests/dWebTests/ModerationToolsTests.cpp @@ -43,7 +43,10 @@ TEST(ChatFilterWordsTest, FilterWord) { EXPECT_EQ(ModerationTools::FilterWord(" Hello! "), "hello"); EXPECT_EQ(ModerationTools::FilterWord("W.o,r;d?"), "word"); EXPECT_FALSE(ModerationTools::FilterWord("")); - EXPECT_FALSE(ModerationTools::FilterWord("two words")); + // Phrases: words normalized and joined by one space + EXPECT_EQ(ModerationTools::FilterWord(" Two Words! "), "two words"); + EXPECT_TRUE(ModerationTools::IsPhrase("two words")); + EXPECT_FALSE(ModerationTools::IsPhrase("word")); EXPECT_FALSE(ModerationTools::FilterWord("!!!")); EXPECT_FALSE(ModerationTools::FilterWord(std::string(65, 'a'))); } @@ -53,6 +56,11 @@ TEST(ChatFilterWordsTest, HasFilterWord) { EXPECT_TRUE(ModerationTools::HasFilterWord("bad", "bad")); EXPECT_FALSE(ModerationTools::HasFilterWord("badger badminton", "bad")); EXPECT_FALSE(ModerationTools::HasFilterWord("", "bad")); + // Phrases: the words in a row, whatever the spaces and punctuation between them + EXPECT_TRUE(ModerationTools::HasFilterWord("well, Bad Phrase!", "bad phrase")); + EXPECT_TRUE(ModerationTools::HasFilterWord("bad ... phrase", "bad phrase")); + EXPECT_FALSE(ModerationTools::HasFilterWord("bad other phrase", "bad phrase")); + EXPECT_FALSE(ModerationTools::HasFilterWord("phrase bad", "bad phrase")); } TEST(ChatFilterWordsTest, FileWords) { @@ -61,27 +69,6 @@ TEST(ChatFilterWordsTest, FileWords) { ASSERT_TRUE(ModerationTools::FileWords("").empty()); } -TEST(ChatFilterWordsTest, DcfHashes) { - const std::vector hashes{ ModerationTools::WordHash("badword"), 42 }; - std::string bytes(sizeof(dChatFilterDCF::fileHeader) + sizeof(size_t) * (hashes.size() + 1), '\0'); - const dChatFilterDCF::fileHeader header{ dChatFilterDCF::header, dChatFilterDCF::formatVersion }; - const size_t count = hashes.size(); - std::memcpy(bytes.data(), &header, sizeof(header)); - std::memcpy(bytes.data() + sizeof(header), &count, sizeof(count)); - std::memcpy(bytes.data() + sizeof(header) + sizeof(count), hashes.data(), sizeof(size_t) * count); - ASSERT_EQ(ModerationTools::DcfHashes(bytes), hashes); - - // Wrong header, other version, or fewer hashes than it says - auto wrong = bytes; - wrong[0] = 'X'; - ASSERT_FALSE(ModerationTools::DcfHashes(wrong).has_value()); - auto version = bytes; - version[sizeof(uint32_t)] = 9; - ASSERT_FALSE(ModerationTools::DcfHashes(version).has_value()); - ASSERT_FALSE(ModerationTools::DcfHashes(bytes.substr(0, bytes.size() - sizeof(size_t) * 2)).has_value()); - ASSERT_FALSE(ModerationTools::DcfHashes("DCFB").has_value()); -} - namespace { ModerationTools::WordSources Sources(bool blockFileLoaded = true) { ModerationTools::WordSources sources; @@ -121,3 +108,22 @@ TEST(ChatFilterWordsTest, ExplainFreeChat) { EXPECT_EQ(Reasons(ModerationTools::ExplainMessage("zzz darn", false, Sources(false))), (std::vector{ "x no_block_file", "x no_block_file" })); } + +TEST(ChatFilterWordsTest, ExplainPhrases) { + auto sources = Sources(); + sources.dashboard = [](const std::string& w) -> std::optional { + if (w == "no way") return false; + return std::nullopt; + }; + sources.allowFile = [](const std::string& w) { return w == "hello" || w == "no" || w == "way" || w == "rude"; }; + sources.blockFile = [](const std::string& w) { return w == "very rude"; }; + sources.maxWords = 2; + // A phrase blocked here stops each of its words (and the empty piece between two spaces inside it), in normal chat too + auto verdicts = ModerationTools::ExplainMessage("hello No way!", true, sources); + EXPECT_EQ(Reasons(verdicts), (std::vector{ "ok allow_file", "x blocked_here", "x blocked_here", "x blocked_here" })); + EXPECT_EQ(verdicts[1].phrase, "no way"); + EXPECT_EQ(verdicts[3].text, "way!"); + // The block file's phrases only in free chat + EXPECT_EQ(Reasons(ModerationTools::ExplainMessage("very rude", false, sources)), (std::vector{ "x block_file", "x block_file" })); + EXPECT_EQ(Reasons(ModerationTools::ExplainMessage("rude very", false, sources)), (std::vector{ "ok not_in_block_file", "ok not_in_block_file" })); +} From b279d187ca6e706704fed73b4d8dcc604142d54d Mon Sep 17 00:00:00 2001 From: Aaron Kimbrell Date: Wed, 30 Sep 2026 08:05:00 -0500 Subject: [PATCH 2/4] fix(chat-filter): shipped blocklist.dcf in the portable format The shipped list's hashes are 64-bit FNV-1a (checked against a sample of its words), so it is written again as version 3: same hashes, duplicates dropped, sorted. Nothing else changes; it can be replaced by building from blocklist.txt. issue 215 Co-Authored-By: Claude Opus 5.5 --- resources/blocklist.dcf | Bin 16160 -> 12876 bytes tests/dCommonTests/ChatFilterCoreTests.cpp | 9 +++++++++ 2 files changed, 9 insertions(+) diff --git a/resources/blocklist.dcf b/resources/blocklist.dcf index ca4db242aa20dbc20082100845e658ab7d958bb6..600443337b2f26819bb234bdefce0c48ed592b73 100644 GIT binary patch literal 12876 zcmWOBQ*b3p5C-5F8#^Z(+fFvNH@10Wb7R}Kv9WF2HaB)o?0crF=d1p!`njuS#6%@T zAi==Ez(FMm1G+#5KK+vQjZi&0SjQLX+QE!cu`vEKeU z+WHUbLoj0Yv+E1Wwe*^Gf)pNl8s?pKf)fWifga_^$s-lIDM{|DYBv{J|0gag>v$)$ z#T?FiRazgk=}$K5o{>K2Ptw&;2~9E>-;i#0r$Gi7Gd8D#-ZCbbS{0C;OfdbEJ{WNt z+AxW<3V6RKbznSleZLk+gJFunzcD>T3Sbnw>8qZIKVe*(mgPtZpkYhAC>2&y@nA3f zUX6UKm|;EmIWQK!Sz!IhUb2$|ykQS4M}K*>Rlvd>F=Vvef5Ns?z*eSvu)<~HN-<;u zQ{l?m^_yX}I^art?3q^I$Kfu@DAum%SKy{EW13+V@8Ir>&s(hY5a568VQblkV8W+m z=H}^rRKia_3s8$}cEHC{c5YGb9>6cDL|=!Oox)GLXq%w5zQF&;6Tp2brTXFGZfr1x zDgOhlF2-OAEa1mGLA=2fN$?LqJS_E|ddLq--a><^kMtiVF2x2@Us*q-ZgQJuw9yd| zZ@;E6k$xheiYgn3k})E9pOYC*4YMF%KcZ1j+v_1v+{*(J9@`L%E42U#s3QoEH$e)g z?(+z6m{q_2t3E`qh)TB|U|>MR9dEZCAW=h1NyOK!*>OS4BVLeqFuggm5+$?SjT{Tpk~5|b<2)y_q6wA$sK^q{$CE#z-$n5 zi#K}m*kc}YfsDBNjCwILO!UC@@oo6{d@E6y_VUDx+*BJwKoSNIRPV4_pxD$$zj zN&|fq+N~O=1_ehHwX&y6i9~-Cl*Gaf#id4+CoQDY+$c0uJpU1D4+I?4gZ zaEmvw6+L)qgqXVj!nFr;=rH}vFbES!xG{S#(Fq`5A28|lW%!nJeXuZ9 zx%3fv(y(Z)`rogWE3p_XWJb=uMzLDfWGSENl(2DsemSA+UL0!8lA>({BB}71Z zvEUf1s39Z{v*Umb*U~?`D&tH|g7i?v`Sm^gfq1QjV~#iSf%v6`W6eg5tD2yL)2XCQ zrY)U}(<511&9CqC)#B9Q8POPF|Cd#dH>7a@@Zf>}xu2*Uj6I(B(}rETy0q~4 z=XiRE4J;7}ej)MeVG)8n{uq8Ih7;8)z8I#8QQqZ0e2v*>PqgPld`4Aj1@&?k0#_$D zR-hpt!M-Pi-##8c0nIzp(KVGKfuatR^xkbKfr%W4EqPECK`mAd%_#O9!4QD8vLAJg zVBXKZ0ap8(AXIeZzVGsZAc?qi4f5!XKmam?@g|*vPy-HR76+jk0AvXVp{l-5X7;x? z;bZwzJ3&nxVdt#3h*=VR_dKt7NTljwDFA!gB>G;Fz1!WvBtk`$JPIMhBoa#w^(rYiq>ERBC_ca# zQu05)^COj;NYPuFuraMiNe7uFGLf<{$xxIvY`Y%~$dpn~?0ajI$cVYF$A@J&$+2&< zC}H;n$*uZ7#)mZ|$(4h-@msj0$h$K0e_GH*kpFb4m)*mOB7g0EFQ~TqOD?k2ugOw+ zMXukZt>_(qL4j`#&*~;jMlmK3ey~qPK`|tre6SD8MuDb*!g^CDO0jh4vd-h^O7XT% zRg03MMX9m5u%LEtPALS{la?(MNx4D_QaO^+-GXN1AJJdRtI2bcki`{B?EZAf$9QO} z96te1dOLfnQ4PNQfK5j#!&JQ);3f<;)u@tWjQLr=P1LgmII0rynAeUWbs%qz|8&X1%LDqW38L zT|e)2MW2O>y@Y#9%fO{5yo9@@$B^j?5~#-@Tnfop#m&FLsfWkd10=Nk1(n0-7W-?&4ZDM}(B5N4{&13!xI8MF z+yj@1f;C7=J}2( zASG3w+VhPGE@F}u$4P)WtB``j#~}`J^6m1@`Of7qKmKkx16VVdF5fo(rZZ ziu5|O>)r9*72_3i;zeW=E#*7&L90bBuMQFm8hi0X;fMkY`^58v^|Bw!*FU+jn}SW2 zJ162MTy{j(Hz99+a%3jfwPX^}Ya1n28G6@ZFo$qfPf3ve;jAyc%}fAfOtyg_FmA79pMd% zE$pfA9S=7qz3j4bzJ-IqhwS3-319LgN9;oZtMx%qNF3fer(AzpC^)cWgX6yns5!0| z=c4#cl{kz^VV7n`OgWg)2$yEI{WZ=#LUzS|C6j0EiXW-n-Uc`d6?U3A`aq1#O4A9*4ZivkK z)4#bF^snPCpNV-SADOXtE68}RZV(%6+^Klj6NgVJURilay{N!rv-o&wF6fO}7z}w# zN*7wE6>@nR{@k=qi}dki*-7fUC(QDo5(!aj`9JggX`|b9;e6p~GrC&E`ugO-3j>My z#Uq0IvWkU^z-z1cwTeZB&b#u1PJ;+jo%b#O3z02IiIAmiPH|HI!rYbhoZ zJ;7PFmc@!Xe?gGNz3i;1vdWYTeDcw zQc))@O`8~2>#;zO6PI{ct>UEKZxwNvX(8?mptiV^qY)3l#ZsJ}+MNer6)ZmKJ+qH% z_bzUqt=l_~MIph;`niw$lU0JZ7o;7lgd>YDp>BJ*1g|v6(Q*l2a86^o$`uLlD&Wc{acCcLy$JN5_r~m2e_TL5_iBBM;fYE3E$M_16%_H$p9<6c(pz%NoD~P zq@+?=NiPMI-&V4_lJX5*oS1TuQe;ajp8{d9Qo3f}zNYSoQd_nAsvpz8q!!;TlB7*7 zr8alp*c|fIr6=0NS7cN&r4jGJ zp!PSmrIpM$3Ya6nWq7QSR5V3)Peks}P$NUR2gR|O z@dqP0qr+;Y6^d-RB{uVCu9Px4e=K_z0{;p*$D-lXhP7ik*a48|$8uGNwC~HA=W+#3 zvSubxdh#^J!@jPEb@Jk#d4sVqjq;iEap7ng5AqiOQd{A&X%tw`FmjoPSrm+OSdIJrv~8P_iLU0u%_gd43C*ge%;0az0j@t}DEV zs+sQXyen8b*XNS+f-6cPrd~nZA}E$NWoe!JqAHdXVWxX0?pON#W_MmLfv2pbdvZh7il>Y> z^2x&SAfoKl%BRsuu%ql&ir|bBe5c$xkGkpq5v7ujn#8svU!YR`3#4g*$_p=TGfQEW zioBXewFRoYWz-~VQmRW<2@BV2SXRma>` zyxN>oRLNi07tMy!Rl69&B3$trReuqqg-VAEsp<+j?OJe8Lt3fzH?Of&(sD~Q%hyuFU)x!(8*|z0>sVg6( zzap1#srM8chykX6>MM6BDYQ;X8VJuhq8UP}8a)m2j)7`A8V0TPEvJwD8nh#t&os`D z8fhzT71sy=O@E{xISp^Ln!bb}q3AT@*%@*gkX1E@P-Su&;M_F-QY(UTA59)jkiOmWQ@pXtjx+X0YwGe9S62>UR8Pgc*vhkcr$`%{x!maHQ zmb+N!P8=T)mg`yTCSR0HozB$j<{(v44_933LYhbyUthlHN+4!Vu#iCL8A`8-!{MXp z5z7hcX9;uZVKMxb7D&9-yYtP5MDD!TYYM5;EWw4+--{KJ0xX*8vzUGcS9jUyH(Zd> zKc)HT`vO5``RHR5`D5kJ1n47)MaGhuhvgE z4g}ex?*bVl1%PVShtwJrEH!`C_ZD9fEVV~9u*c)7G5LXGprlEmn&ru1pe$j#1SHos zNE9iJCu}t|urG660@D9y(2b}8QImoDe1Mca)B_QCy zz+?LrG!JG7%?Yy%WJfkU7Y50TY&a_iQUuvh6-m4OB7?|~U>NQYnMK49{B`sz@Y&dq zLMR6wORdqc3|X3IL6g=|i&A}6qEwLy7dHpGzgn+}>}_OWo;vJ*aA`!wO&vo2S!+rD4J5Ps?}=wL3oNPm zzqm}(;tb6D{{;R7ytw>^F?IhTqoN}fWQt4!Qai}Bonf1Xbhy@Zsv$EF%VpBEx?jGf zCu-g_J5BC=U3ABET+rn_yl>x>u!kYta0blmT&Sy)2IGer<>Rp%g_)umR_c&Tt!tDS z9BlGjW=Nb_TeWH{)H}R+Kq`_h=dPrAhZ|>jrLLYiT~jAq8CwfliN1rV6j&3A6D&QSMh?f+?Rk_?&L&Glt&Fk!wvivVG9bBU7NErn=- z`XV7w+>c}N;QSwIB}=cxmH)gvfC0)f(ERDYl}35X-T#yjMgBBcZpEmPa%cf9$L!Ow z_A`mBcq@K@#pRP&4J;!+<6=Zvg_t}9;q>QNUAQLJaoCnwZRSfooFMPDF7-m`ty{t4!Kbz=)dJkE4A ziDRdd8@vuDWocJVq{6Fuc3_9>VzMAtaBYXt*D)}di(o$;sQH=R%WuEtUlbCYpk;5U z6vO#+80-Msu#|+H z(e4<7T{+B}fa5f1!eug5LE=P^%x+-=X3YvR1=?O_!3VD5~6PtKl%8RERwhqU4- zTkf2__#p>B+T|SG=MaoGG~=ApjGa>iwe1Y`cT4|<@!eU^#6gA$6TzieNzIdNm(j&E z$;dW~xyVJENRb=p0PU*smy@Ei56;z+cfsI>9N#s#6caY;?}@8xPl!F9w6vR%K4bVX zrN8j*{549p*VP6c&WNx`2&MFHWi+Oj_e z0|rt%7iE8p%B<7@0bPGSVja{0hJ$|&*5q({rnmlhExqss-{QL?OQoK7(zCi}RGZnY z!aKOvn#bvluH3uR#L>QFwSjxwI`t_dsmXct3FsIhVmf%7+az=b);W3T(4mc6rnq`E zlCG>RLr!=!6JW4_`%!!54Cs>eHL!RNi5U4|A{lwU>=_rTOE`GCoAIOkR-f~H+ZROn z?YHR3h`vC1TfOM{-{&IBLywBrl^R)fe$5}R*Lzh#ciR>(z8v?O zCOc?Tc02y^Cu#kbpD+S6F7>=4*8AaAgy;~Xf? zJZ_NZ6dtHD@Vfw!IWG|UXF8Fb@K|7p*x(#B*nA)+Y*)3M{c~U$tnuzX+Dl-!+*2EN z(_3INZSL+qxnPi0fgdw^pH7h3kkF>5i*eAI{{r}GUw=?UU#47i?{JV#v|C35CR%V< zMdze~8eVWBt;>q2fJ!j5Sl_(1MO<)>!TU!ALr(Aj3v5M!a#wK0@&dN2_H8hjxoTt7 z0!c_?H>c=YCtZk2Ccvp%kw3(&wUpUh-6-TQBt&?nsA3p&jW-hQ(RWp+!j_d{D-)VL8g{jK8k{VS)M}9{^!%Lnq5)!!%)N z7Ec)Au)o8oeL-UW4kMaUDwOus4C5y3GJlHf3EPQ7gHUT+3A=wsv&@sj37<>v@*@#p z52tS4D;P4CM1IO`$(kutMkX5bvjV~PBjI;# z8TdcXBLyI}p`5&bL=6vHK{-VcM!BICf>}=!MkzVdkj>dsMq!v`ZAOaGMzIk$Y)0~l zNA29MwKtraMj4Os!2>#dqLlH}-~mM$QRh@3D>I_(R#U`vkw>B~eF176AE!|>Q9Ubh znV>g=k3ACt=>i{XMffNGGnFjXBgtY z!F5v*@n1l^P_gk`*hNIVPjHF>+d@`6e~rgn77< z^Iz!s*op9Mit~8%f{Ddh)0M3oB8f1#YAL%*>WTaC!*ad7`iXa*b%Av{DTzc~T52kc z8Hqs~d)+Eu`H3I@cpmMLsuI8dgI1TS2Tx)yY9sSJ&`-*k$gb*sGD;dK9Lk{Fbx(q0 zU9IWfnM}%2Ai-4e+Do!j2kEz$w3kk0L+LM%ha7YZfq=ZfzrmGK1X)|$@|RDRojqu3zvAzg{j3?k7o_ZYd{uIkLRh* z!<*!~^hAKlS1h_n?Zn~EN63j0VO|%?7tN_VbHA3$KVp}<_t`Mak97!+MhB11M>wX7 zmVS!Kw=eYate(iv52+R8*f(m-UlvU%U(#yMr?Y8Ym*$B)9T zVn=l0-3bkd5@b;af?q|sB^%XKN860WB`JF`fc4Y+l2^(|@1bt6ziR6m#oF^ee~}%* zN?LM#|MGV*i|o(@|CJJN&kbCj`a3xl>_R7RTKYvHeCvr`TFNJuEXwKKRqFD##I@f( zQ!2VYIgMFES+;;zXE-aMR#x0~4|r3=Do<1g%7|iPmnVNRE@0C0l=I^s|NVK9T27qE zm_ar&SN>lv%B?$ieZ}JVe^i;;BNfFW#rZ+A2$js;=Kw?zv`YPW{@s;9j!GBP>|moH zu}YO153>k?tV#*-lIP6Y;z|f-D2L)iw5kvI%i+s+%c_3#dE1xWg1sJtk<9k1E!RAQ_fY_-X=tAlkVGQgyJg!sO2%q}72Nl9nx}^wrO56@T1D zHL8nUFn?9*7*&Vh;zf311y#%KF)WNNRaIMypr%K6OjS>q)T50gA64I*gvVgnKU8xO zVNX%e2-IZ#3iL;9EK(6eUwC>%jk>axb`Z*7a!?pw{z$qWL37I^KeCXrdR z8(OVI4sc>xg1lBhTWPSigSs}%6XY~??eCxbcXHK{wIB9D30xBswJ+5*QC02Ib$36{ zRHhH*>b@9HW;aqi>+1US!Z3T2>*7vr!;x?c>f)?g(yptP>(pF{Z7_2;>eM0z(ylZA z)%n2C*=#W!_Z^VyF_Z%0VYMiLpu{gB}Z@ROTS;lZy{OHp1>t9=#7^Ptv(UgYGImC@E1q-5qB z;jvcFeCn20WRf;gRH#7GSkE@BBFsS2fY>%%Lg@MF(e}0&)`Vf##s0Rlz{gIBS74hm z*Vjo`_DdV`HKzDVEpGcGGWGfivqL+v;b2F@j88l6{T;d+|3P~t+(~-OtVl;c&R2TO zt6qmiffxi~&q_zI%zJ#d+2~V4ecO=&d+vL?8Psd(s_lRNWhi^qvD9A^Kq1ytXfrSt zz0-Om+%>=jkkv)BxETmC7Q=H=h99JoGu1^5M;gTbk_i>(ryUIV?Q9ug#6I|I%z{AG zR%VcPE}C8Qr{SPx#D_=yzRO@7j1rp0Uc}(!sr&62Zp$FQNLWA_V(TEZysHxj_~2lM zt(NLd$JSuv%NgTAJoL~KV`KJVJ@U|ju}ked<9J5gKP{MTEWAgus>M}z<}^oTz_5Ozb$g7C!C!_l!V8VP!17T!l_`wjJ3xgo ze(H>w`+8!j2M<zDN;oJgl={140jwbHcE=X#8^oMQp~mM1K+;2vd-W{7$l)-I;~BsWbS3eQuLcg3 zoQT(q_fxsCp;WJoCw%j0m}l*eU-wCCn2!U;U3?8S%mL6QyJ_PC$*Q8Nd?%ui+ zZNHS=`?O6a!Zqz4aCyomObUMV69~;r@JBWSOT7P0oC1}xG@!{ROGqx+AX$Vbm*>}3 z0?dOZt5iaZ$T!0#r~kR#q?O`NSz;7v{A3}Y@=nZ({$Z^#bq4m0ijiwEwSjKQX*Ogv z6@dMf_o0F`ouW`Rl66Kmeg9EAk~P9KtpyZdLKrQao zCKu3L1poBT3{4RTY(lr{*lo#$7sv5(7S~+4@>0feSS+ZzkCrazWq;U zK4F8Z#9_~5zV=J!7o;`w0yzM<<;AYNK(vE<(x)1^u-w9Q(x+dtkezn!ioE&0@Wg_& z@r?t%$iBzD1lb|FSX^VNx+t5xc-{#Rk4k-AoM=w`fZl^&@?p}tyu8C-;b5=i;m zrKK8>`nO9t{atUk8qZ4ze$#KbCJ4*$n7#{~Blyexsr-b#hC$2spWWUZdd1826FP11 zymiZ=cza2L^ykaQ_oK$|D4HvHHtN3yL&8>WaG#CezhYM;A7Hq9coJ8B_f{78LFTUP zb-|jvzZb61Y=cBDTFKx9CjRiKTWRT;-p}=)Te0w&?M!dOT+J{|oBHF)w<^lff9ZDX zxatCNj}<4Fw8~>uf{3|(zZy4JbR^G4u$KNYXF4s-wYD5MZ#tdMwYH1ud!@qOyv8Eu zdMBaPy%zudg~@bxv}VQ(i^a5cvzEwh3wy%|SbqhW@zgg_t=k=^!`>Lku4{6a!rsX0 zt%s|E)YV(ZR+EJs-H2E}@3OW=bsJtkod)SYygmhuJjrZ3xo&A7{w$=tv0jn@g=j;8 zy;1LASz7o^vcZh83RPx8ztNAKQd&r^w&9vR19w5{xuJd9KZou=ztOL12F{qYv=Q*u zf59#Bym93Fm;$)-Zjq6aZk>E;Pl3L_gl&Xyb8^U&s3CLdv+sT9!h#pwz+p2jRd{Djo+YLUp zvBNL~l6+%_D_$8A26|^lF8!NV`bB-$Wbe`q+q!gjCz>3%4mGjszzOp&Nn?5UUWw3#X~xs5nv8c(Ye3WfdP;ha_Abz;ddYDw!yg7It2t$Fu|>OPML2y=YQMlZe5ZS_ zI)mbKl=E`W40YN#eE4=xNKYh5x#VuI(IoK4i7@tl2YAqrlU(e5X`buy)o;puoZ6@I zRdkkpi5MY9|$5~63*H8!!84E2afcQFx*UN(Z z;U*P_3YGeZ+8fu0IA2Pb68}*g?Y|d}80vBy`4)%swLhI58Jc8iHZdX}+q97p=j(7B zqX4G_uMqi;=c&^nNqG2<&oRrP058_Zg5k7Cc{ZNMkJPPDfIjbI?N>8Hha}%)i|My| zYsJiC&JTyv2g`}$a-%%d8TzB+95f|-9W9QN(`_wiKsoP;+|3vS`X7OlIzSe_T9xRD z_KOoVphN5A+djMzPGRuGqnc*1dUfMu{a9(F2cF~fat_tAZN=!cWk70b>A$K|CsD|= zqLIeaYcdMuhfw?%ec_rJPgjBfVxkot@w8$w-cOUYt`dlYE_B z!dy&-FNL9)` z-w8q%MY~i5;@F^%P+r!kFwdE98eJOVp-Z$PxL&R_;lL`g&R#(Ej-s$bRrr+O4wWpP!|ar*ju&-H3;BRkA1-~H<7 zU!c&NOTty^-ns3TWZ4y*gy5J@YTH$>yN>NvP{)-Uyho^5{rMH-8!cTi24pus z(278^+}Kfov}UOr2BzFl+vE=tu;$!I zc$=q3AvNESGI-GECf>|mq_09WVT_f0OLpcxU?!re5He(Nq-`Do$7XLkFba^G1JX$`cJNF|3 zuG=4tqdw{wsM(+BIiq?rC^Bt04TIPTqZ*a{vO!_K2TtQy-de;n<&H z9C4a)Z*-r7B0+xXK6?Sl&-}fYo)IAQ{}~Y^ylygnP}i{DSw zzsf{+q0tbxy&kiFASa|uy|z0t|3=lAe`QbtTpH$WziQbHz?z|czm{Ir0&ak0Z#h(5 zfSX&Ew|p#enh_@1x7VD94GUTOH}KI49L49Hw@@pq(#y7|H=?@FRXBN=_b2~GmGMT( z_l3~U%6VVrckwUsmePKH5of)f8pzA yvbB9De63s#7czAoe3f0-9b|2}e~$*gA7tIkeOEQS)%J+be-|hW%<9yyN?~ z2V7&qa*1&A%=^BkxR|8qI|vAfZ8-1;@SkZ4keZ>b_c_~TRuxi_ab2-DmEvLwKAZ&h zcctkS!LNzn>{0Z49O-$x05+`i#zV*AM)$b6qifT zBO8CP+wbsVGM>|#MG6%d7wK)-wHks;T2kIV!)Qp?LMEDhV5jLE?1n+zk1xmi#msS7 zM)WqGv~kN+>e&$7zLgRx#bsh*Z0AhVup$s5D;yDH%5bOs5pk_Y4@`E%wV-7wtHuL{ z%IEZKjz$7|McMHn!4M{xufpmRW2>jV{VE|m1dZI%6gv0LFrZfL{RTOFqM%quZGi1e z;a}bG>(xSJpW-^C4qGV-Z@1`1tH&+hj&%qhQ^>XmkTzg5P#lH7euu! zqoJg$h_U_UDqlhq#xjN#mwCVS*!Um!^VAZ8G5B4^^yaG`d2u6RyifVvM+jN&&FS|h#1px)wriK<^prJP(l&AwT- z&EzC`x}19*^0&=6KOckOoW2ALUQGM2bwp}+xb11@Phq?#U|h}KLKu+m>M}^c5u2n$ z|BT?HvYqyTQp~Mk@Bt5v?N=MjVxU*syQRC>;pp|IU6JCpV!z+m$ z%l!23g+Pno(V4K9@@YSor)`?}!tX@Pbm4aIDU}Dt({JJ=*@e-BZ8Pws8M7`^;OR5t zlZv!$ET-hvlN93V8GiPwlx4p2%D8auRzX&m_hjsUmt9D)yq=x}Cj9dH0*o-q`}D@C zWSrUVW+}g>nSe{(pnlRu4KEQ)5UChMQwdCZrh$v^vqNw^_sB@~%Qp>CVn!IAo)iHh zKcnFJ1Ug0hPoui9y}>H;Mv%DdvH6qK+!`xauDWN0-0xhs&t|8h1k98Jve zoj(qdSo@gsRlJuKt}I`rJ*2~1vN8Fwz#zK;C%X`O%_WYbU9?Ov508`n zoIbkz2V{7)(y0C?Rf%*vZME~()xMNJW5K~J89W)s*gW*BrfHBM#n_pJ*= zHe43_{29UQZtd;ML3DUL@6&H=ukHa@gxoUb+qk1UL*yspeZl`a#SG>n|LenF>n&f= zasOae@chL8YE&N~<9#dhy2~Mb<-2T@_mF+PY$@mPNOZ0Cr;x?pW=tP9`^2hh`=M+q z0QoXa`DbX2Rxuv5!Qhcw>#3NR;b|#JOtDCbs+ocgd;o}l@HEOapRb`JtV`tQ~ znKysY5UN3@vRoRRxrIFhEj-&|Rmb{hBIHmRuF_rhW4DyzcZA|HNPzWGW6IUlm1h9Y zU}Z5uqVf>80rjfC>ZzXAFMhYw9hgPEyE$yd@gD3CeC#>qk-S^xvdYe|k))oZ@w>&0 z!1@%9ZV4vbr*MCCt;f-H_ztJSft0%6d-8C-U>@zs6WAcIxr>lfCM#A9W**XyUp-e3 zW)3os?}dITKQ}<;4?upOUH3?*wYAI#ck@#w9y$e$eT)eILP6M8yeYLvwS|jR|8abTFQq&E0YvoSQ zk?L$csCkdtKoU_Y53?QVG3Ct$Q_(D#i_kwz8I~_+m}wR1X9l2{-(UQ;Z=w&xdCh%R zMIOgCnVUhex;UenCw)CUfjv&NX59{Ea!*Swpqc!{Mp1CfJCq5L*zozyA)+2$v2TBW zOF6n^tHbAVs4^u1<9okUwhkufqcH@`XVxUCiLA590W*~*n+@4e)9YZ|o^0#VfIE4H zP6ZpXV*aCL=OFi?kf-g;^QpH{R~=LIrhA0H&LltaIFQ&lVYO{1V;={!E`A+2w{P75 zjJ!W&+Kz+i|8Z8cow@4q7Q9<~vZC-w0ANUnrww_-UTpn4=Hn{F#ki})qmLMQI0-mC zX;3$DFw`}mOak=^v3(3|v-1oWKbdUga+-m%6TL)|lHX9fkAw>7s98|^fr?mE1Ndgt zh4*;1gSfNlcVF{c4<0^kx}C-^Qx`jKnojdeCr`$|!;@$3ML5aACH5;d&PAJG*xmE{ zJ5^5WHS?*ygG=mp8+~XDo$NteD0PFf-u^jt6TS*-yx*qRDKF{x^d>V5>2t;GX&<(TdwNY`EDb(cO{Hxfr>UQPzk?6rTx=RF z^Tpp0buP4GGk%v5OgPgQ{1SNY`U%#IZ*cCtYmkA`aNJ*8<->)>ls;%mcL6#R;&r_J zZnel6rfKjjL1!}G$h_ZeP(0Hphz9E}XTk8bJ3~IYYakflM_;n<#5-P9(C_?8K?4R2 zQx!z8?lK+>4_d_L(eCKgnKp%t8MI9F0QWeaQR0)w9#(ujjf1Z(T>Q;3iQ-9lurj8< zMWCl&E9l$(RT0#G@-NYHCX8qG#-wkZq@p0ep#!-74gH;Z{X)2ws^_3u3nFnHil9au zZ)GlQHGxJPTWBu-g&aaMa#i#cWb3$~a+MZ=K?ZyCTl_1)XM~f1lP9YR0fN2nAC6N> zx4Bt$b6@gu2G&Zk<7cCw^`vVszSe7S8CRdvLK z0S;WpsoK!RJy&0yq@Fg6qv%GT^yZk;?zBsRSJeRQ-**>Kt8ELyUeGA_l~qv{EiyIk z0tJ=w*IRWz+w=Q!$JiMHQ;_wzHl!?s9DA+E?0FilxVUaA&6X4^rDK?7RBh`UX_OcsV`|3=}h~6H#T63*yA#HkI`_E+cR)RhKEa<$_V{t4zLIJH7({r zzg1GA_%5Tms2sz-`Hfjnb_% z)B#^c#@(E1Wxe0ffysQItu!HYgr#5D~k;Z=qeg3eM1%^lsTV} z6A1LG)?JYRl9<R1;{Gz_E%E7<9Ezg@b?J^a~rq()|Yc{9UnOU(wBo~8UL8~ zio_nQ&8Ko%z{u+THPEvyI66*M663ZXYH*N(D5tonm-NSZ0;6*QXTXt+w6aQ8z81W7 zbH21-05?)-DE;r2wl#rDv$!;u9wg9N%Z=0x3(iIvIWNF>$p@uYM`G~+ZM3Dn2a(%< z2Iyb^>CiKYEg4{A=lE7mz|?|YXa|RLA0sTl1TiSvKD7*g)XOoUQzZVqyN^b4r>%26 z{YIz9y&P*m#4D9tz;0&WXYX=J&!Gt z_6gjNoVSa}>Ei0Hh;z1+<9)MGe=uR^4z#w9Q}3ph0(fw#f~^TdG5QTw8C`SpYahVI zQOu>uWTHbJA##jPf4TS^4Si+^mgxf(;KTB*wPX+&08OI6zn|xp8^|wA zya?8rGytEm%!FNhksGDsweEA3&Y(PV{9b3sLSO+-_11|yGeG|WI8zIr$)JBF&K0)N z(Kd^R2kHgT8OyB##330Ywv<+9h>oIvBCxslN#O#M=iaA1BkzMH{a1=s$l&+;IXGMO zra|xbMR2yDOD0$>G-E?KzfqlQK81EI?I{m-B5)oMemiEd(5eaD8Q7fb^n-lYtg$fH z8G8-MPCEhrbZFsY-c-V3pCS3e_7i$3HLKS(YU`tk+I3HvV+^b3`k0Jne3 z<93d(egTMYtMRcfG<=~yaIrQx3YpW93-(W<=KKcbcLBY0$hTF$gHk8Fhix-cQaPyc;d)9jr!9AQJE7ni*7xvu+tpi(i+&5bv_Z z(uLJD4+3m07wot&Px_^f1jIl!!=0N7?${`po+`Ku89Ke`pA$ZMlS}rd zyLx>*^&YmNl>WRT@f826K)TBTFha}z8jttK$n4vhgN5bin`jgNfd}hKwAC1OGEVJF ztLddJ{%YJBe9sKrV#E&T)vQ$crzNp0{i}g{Y@hx;ls= zXrB=t)*wmZ_je#rF#9#PL7uz_ALe%>5VfRFs`TI%Vzi=!Qr?Me=~=**#9TANJ!hliE`_nS#pk2dtYx8wq~k_cuVwWF zr)#dv%&1>kP%#+D^3Uh^VyhwRv>c}s8|7hB&<--oJs#zJs*Sd>k9=U`J~M$?2$kPt zoEe_0s4Rkf!}1i(hxK^~>9{mv{-{!OOnl_`(}F}^$RJaMTmPLi^!iB-p+=C!?dOGh z1uNr$bvKe7%=j}2-L_l06cvb1rqyMr>cJmC_T{|=9u)^pR z(Sox8{SeTcyz@f8+adV$Xn#~HiLV0C?eE^^Dbpznz`Sng^mp^Mxjss(kxZ2kW@vbRlCs&AtQP!}oI+i*kU6#n-nQK_q?UcO@Q1TSd) z&A>adYpo-Mv|CWu7zxYKlX08{>LrS=u%@ZQ=0tEHR$aNvP<+Lx&2;m`9ml3*L)n}& zpXDp0ti@EISR?!i?V@2dm}GQx2Pl4@Csfqx$rgUkLydS6Tv2c&VuiZ=N(;4xBPJ!- zwi^2A-CPDoxrO|YetURSHJgItX-IzF0lE2#=i9QHK9W_@aci6?G2bj%Cf);SPg#iL zX~TMceuA3RA&q27h6Wu(7VFh)ngo^9!ToB1U5FJ#78u)1d3l-G!GB~y$(AOvJHpLO zHau-c7JU~gg~t7*r;P|pl7{q4&krapR25CTjvGVr_vqmFHDo;}xf)DPv!HVjYm_5iaa^`#CR9MhF9EH%1S|MoUNg7S{Pp?%D+TyF5{*huf;3 ziKZ9FQ%~bOyH}6AyN=bXSiEnF24Xshs7HX*Gu|1}YSKDQ*JftUSq$p1aIpSBNtan> z^CFB8KVXo?6A>3J2GkNwK<;06FQ#s`x&Ea1!L?Ry6E}*Yd9OZ7X$IFhyHkaFjdhte z>4j+qi-EV)F1nG@c)POQ0=v0@tk+5BIhEK|A_n%$A;cRRW{&09-^)fPTIvVMFiCV+ zj*n5RpL0+zJRob!TDiwA=!eD{t9EkdPNnh!WtKd^o>5=SZ<;!=W!y^H{^)Y?TpH`aC}QX*S*b^EC@W7haJ2+srL~i| z+6_PFu{2ybLZ(N%5iol&0IC`IW@Nqj5J?HZBYEOotp1?^6X{}={b+a6x}EUlz=G$_ zNhB#WiHH)(6B}bK*hC!PSp}i55v92Eh-hN{qJ5JRgY)?UHoEt0Z=(~1=P5-%d&0wORWx z)bs~>wo#{QK_qPx;AiT^4@3E@(0jpNsSByO!XJk%3DxXm`xxAjSPaHL@$hwdPK6T` z$4x|aE_|c-8zoQ$o-vB!BANOGhejAgSvrnDNTTTFVw^oE?iWsrMat@4S&Ma z(hzsuBneFE`8uyoel3KyA7g&%*7 z0|@gWjaXh}I0&Fp;c-eEE>B*z{S3koEdI{1TS5(|-4AOx=Mm%~fzo3>8Swvqs8q6U zQcSjq-Zp198Y zO&o29P%MO-{Ox}0DR=+18}RSnGj7#>F{ zy<>*?4XQDG0rvO=(Wtoi5SQxr^C%HL_&^YSW9*pU;1h(>fp+$JABC!z-G2Czvj%Xk z3jDqX7jrdGjhDWJ8Y9J|C+twHBBRx>x*wrzSZ2ltBLX|xM<+MX{7zLl8h*Pa&J;Eh z>FD(#M|4J1W=9MRSVKEU5k*NOrJTOIKvcqR+7s+_=9cIul398uoxvc?b+j8dv1iOX zYSuird0?|?TK0&iu|21YS^(znxqCQ%YUdw{!zU`z)Xp{f!(_D=E^EV~{N`+7Yd%*! zaxY|yVUk+puzc-jTSChxt zg#NIFagCO2pxW#R`&P)=X6Y*JdCF{DhWqz?(6rfj8u#yKp|6uK?XS=PF=4QfOk9+c z*m)17zKU zEFdZu_8%<<4RG46@E@AI{l;MVgz-nYPrMMJ^Ij=MF1#>GN;I2pm_gWv>>KQ1eq*mv z!lMy$VU2j?$1j8qeq&P%!q{gei(aZ?K914g8J$u1m1%!+IoYdb^*z1f(WpL4fs&v| zd-abm=|lS8dZ8Qw&HIH{yj)F!tN6cec~z9fte{1hjdMdpW-BL|jq{>J&PTd|{~dPW z`vz~9@p$Y)2*+rb@l4u6E6hU853CepwN~YRV^1!^Yb}Aaw{P4HnFW*r-bLChbVg}E z+Bamo4nh|)G`H$?q87GNv_2U1?NcW+u_NHF%F~01h`UNDoQ;63UT-vO zx{f|(zaEf4=ei&!s2WntBA^2xI7knI z`9|M$fI(9_%8BR?ikz~gN zEN(}1Rfadm1t^Gx~ZnYpRozr;=Z11U4gN2A%g zhha{)ZU(=`Xi`x9#|Qa&sl`7R?6mz_q>JfM=VwLpr7^vaCh32TNEd^W&-0c%MbGon zIf)AHJY>ZR(L59Vdt<&j#}@lKFfTs+f`^nHp;#REz9lj0Ypt8`ZIRN*@m+*cFg@f3 z-5;!zLE@c|Fa7XZRNrI1zdLAuhEUv9mxVJ$`*>ViO-JgdML$iG21UxtzjvB-O22qB z09o=gQ~SsdO^IJTNsP;YD=kuq2iY01@=CNf5OdxDjT=@wsaAL zb*n!xH+Qi{a1vm8rMdkgf}_WZs7I^~%g+h>+s#)iovc#HxI1huijO z_mYxDJ-C+66JFkxnR;w06o_I}DoV#y8Sf8l~m)aT`TWx#4I%nfdG zr~vaQxXD$ywjx2g&{Sx5w~MC}&0k)7e!ZYx4a ze*)vgFpDM%q%4TpzP4`1B|>vM8e*;#wP1ziJIk4yM(HcK%0=mNt8Ez?Qi3V4Wq4Q$ zX4ICU*4Jmx+AY^UfR9^3sGt02y*$hXP(Nu*dDVVa+l{QSJoQB2tMxKiZI5hWM-Z>e zv>U!-MiBp%YFCqQJ5iudMQob?zK2kFVuZ*gp9#;;Q}N^*GgRV>NZQBtL7de7>}#fb zD1u9I>QvP{rauMSBx1jpv%j66{32>p9MvT97_err!uk#nj$Bb$VshiJhD6G z1e;m-Lx(fL>J1Y(g88C5yiAN5&V*_8KlFG+0e{|dxOUUo4^_k{eHh^+ zj0r(_Sl<;(3{_@{1t6Vu&(oPoIkv}p@cT*!n^E2Dxl@`enFGw1mC%^bC_)j2?o7=R zj-+ZfK3&9qec{Cq2NUa+*9JHHfG|>wUKHFw^5hEzJ^MxqY#)^$NA#7Zxd|dD_i&xBG9C z#o2ftUTzMg^l*$2cE_&$!`p4*j!oVI@#=dE`qQa4JQCF>>0?`$y!++T-uPthH8xbg zb5XW+!@B`Y5G{!#YBDnX+x~WufU=0}MA1E@sGg@@U(lxP^piLdp-UYFb_m|f`UPwG)Va7)CjvtJ*FZmIt5=;cx zE+Muee|IAdqFQ?JEBZ& z0W>EUdWTX1#f0-J8zvz@*VkvtZ8}y)>OAFm2_oMlxnnB}mSrtLsBHDMO)v1h&(E5D zmPZ(Bcd?e6BL!t);{S2FD!tngDaiyNH^M;KlwYvJ( z-qf>!<-T>iaVL>oawR&R|*TFoIrlY$EcNHqRom z?ezv^hoPUbiUlOXh8?58%TeU#7jSLVsp`FZ$yI*gFKcT2!&(Cr*#1R*S;jp=@+POh zpt1*p@;Hm#HfgVzM%;w$5Hp8mfK}0)1rD2e_563-+bb6H-`(H80}4Y!;(noFy19>b6){Wi13|$BhFL)BM-1Tved#h!lU7JBy!?g>IU6xC&i1!@mjT`^@dxr>hG-ZCH)l@BzCN!!n#Knd(&iu)HQ(vrQV;=` zq6SDW^UKvshYzQI(L70MR5XPr9y3C-(lkDUP{JzgOq_hY?!e-K3@3_zcqsa74tuuX zCL~tpM^2P1=g{qXFjJgE-x($7KKnp7Sh>{ZQ1C%2eZ9A@{2qc*>k-pOheB)M;!?vx zjdCer3xU>_ilh)pzbdeviez3-UqJ?@Mm@bOgIf`VV+5Of38D*b$X^vNDexQ+!RGhe zl4(tj&%q6n(?ha6@1>SE(?g0t?jukzp+TyX$UdszN7aCb%!Ag>Lws-`Q9tgxF3+fb+{BXk}T?oACNij^Gr z{ow*^+1a$HQJ!6G)!j&Jv2vDc0B+&a+hSGi*Z9JMS+HN@`C)s@nP ziqePidas=R5g9PubNY$fwbfOb(i5ir9qQn8&Iq&MFlLBRHtkULB-|xMt4MN5W)vLF z{vu(DEKd_(=GbuxtR0@9ERAW&Ow+hx4sYF;^QLjp zZeG8l2>zDa!`%i#OjbyL8ZU*qb}lS=DCvbGZ@-)^fjaI}=rcEE6kKh4CwOt34(EOl zTX_%KZ^Bfq0V7aJ-Pcew4F}@Iqb@>%OY|d6NZbazSG_x}*dPx|!-5gpThx@4{Sryr zTX5`@mo?Kg182DnnRqyJco%exh}kisEK9<9;bhx-phUhPZY7a9yj6Wh-DvmLIxO^% z3^(w6sn5v9_OgZdflZRH3ru@9$x)vPFrOr_)VK|hcvG?cQX01L3=yU9ezn;3lBy$F zMy1*HGBP6h%5v};O)1K_$9^1i$7yAnZ}*Y6VGXdU*r9ux&p46~`3bq!D!595ZA{>CJ_*o9n=I};m`}=Tc!B9UDZ3F(v7CGOasv>vFaSGmBGBOZXZ%-aZ&ut7-!ifD z{^l7pSE@BZKOtEBgf?nquwpz~+X4R^Z7(EOYiC((U)Tm$za`vBe=fc?_k9ErX5|$W zICxGd?2IyHt_QB;X>baOkS5}Y@*5wH-U-r(uCyX3yU>zPIZ#FQ&ZFGnVMbZcsYahg z{5U``DnppC!MVoERwTtz_1;Xh(g4$cGtqMVvDQW@V-Mg;p7b3a#-z0kHVq3t0aP}? zW4vIa!S0N$Ybcy9xop}cQ72=X(FBCKLTPI<84Y2e$uK_xsLeO3 z&XK1s@={lf;>)MJNR^*tN7U*Fpz`!6c9Oc4NPdYZ6@2hqm>z{`p1Be!P8TIE$1y@v zLi6dDkn2%I_m){j-D7p9#Pr+EX-j*gy`TD9BH@5D3coCV0f|!N0cFaE(R4zPHe@x6 zH;GVn3lxveegs+%37WrT_T|X9Bt*VIo?>I8dgZL5T4SXZBwwJ6mxy*LdMc*+6P@qj6(zm%diJGs>#%iRC)W+7D(YEqOrUf#4U zn(&B4hflZO>r--C4LZLhf830YZI$MqThFF$2c2KAK6d-Q$V>kSB{uYL4_!^oh@njd zvp6G$g!}B*fGihht)x$Oy|X%L@ktX&~-%1UikE@?Ieq4oOQw2R|?^>F{n zorf*mpw=C&%Z=u<-JsQZKt`Ao>xk94bxv5CQhe`JyU|zZ@2dvqOiwOWd^3jU6a;Q1 z=s&mBqO#hEna!hI^t61L?09FoR*|Q5h?RIQ9bw!I{0b{M2rfRtkA?o|%36yMjj`fo z*pJqY3LrjfmC|9zgqlwL#wHlknuYm++BL;cIc49Wt2&XKZnes}PSFW9%tpa%r|AbB zyrj^uZ%Xs2Th;{pF`3tL@bYGd>T6!Y!d8@@g5Db5=m&Y(-aa88jiAtVf|swf&rufF z_L3aF;y1=#X69z)rv*$$;>F;`+8fmRmLb=r{cnUxZ1 z>r+}oRfd|kjDG{feRL1Fzcf;EPNlW`lZtXgs8)HJe-+FuFPopianrd72~Na0@0o|{ z{tAPpcsn|tDZQ0%60zQ~iF1QaGwz^|XpS~+5;1VGS#9x6e`w*#qsEi?-g-p4!<^lD zyXyxt>?FNaOVP#P&yy)MujU2ggOtr9yH%+lr`4+9ueUM7{`Jd)LTS6w=O)G(qoftA zqOr~EBV@}y&!Xo*NKBGX)xH#DMjR>_XFC7V5A(w zb-NSd7{3<;{<-}$=wZOG_qK_&C?&>9mPo@=+Kd&J2F;oqeEJ~pSH3c``k-f*|HUbc zdHf`%0Z!>J{I`#^0ie3OF2p_hXqnrXJj*{RP4HQ?if81MD%T~jhuW1zo**a72C~S5 zDcv01fHsD$R>rDf` z^nZEMfFBKb(f^wd{a+sRzy9;T-ZS7e1Ag*&Mc7SIG_;vr~mjNExzy28Dg#kX-zupz#69HZh z;L8Ah%zybXfCmHkFaPDe{Fm?Y|LUy(p31*I%D)~8;GY1V#{a9I@vn~o_zZx*0Qd)h zcK~<>fIk5E0RQ&;Ko1Y}?f>@aKz|PO(?FjL^ua)n3-qx--wO1pK!5pfZwd5}K<^0j zi$G5Z^n*YT2=rV)Zw2&HKz{`EH$Yzl^e{l*0`x3EuL1NJKyUGHKLPX-K<@u3xQnlf6D=ZoDIm&fV>OH zpMd-b$bo>o$0Q0DC{M-~YQ; z1A8>EHv@Ytutx%W9?!|u&-ix_0G``{=WXD57So|{{UYoH%$Nl diff --git a/tests/dCommonTests/ChatFilterCoreTests.cpp b/tests/dCommonTests/ChatFilterCoreTests.cpp index 50b66590b..b4c7f71b1 100644 --- a/tests/dCommonTests/ChatFilterCoreTests.cpp +++ b/tests/dCommonTests/ChatFilterCoreTests.cpp @@ -147,3 +147,12 @@ TEST(ChatFilterCoreTest, DashboardPhrasesInWhitelistChat) { lists.customAllowed.AddEntry("stranger"); EXPECT_TRUE(CheckMessage("hello stranger", true, lists).empty()); } + +TEST(ChatFilterCoreTest, ShippedBlockListIsPortable) { + const auto parsed = dChatFilterDCF::ReadFile(std::string(DLU_SOURCE_DIR) + "/resources/blocklist.dcf"); + ASSERT_EQ(parsed.status, dChatFilterDCF::eStatus::OK); + Lists lists; + lists.denied = parsed.list; + EXPECT_EQ(CheckMessage("what crap", false, lists), (Spans{ { 5, 4 } })); + EXPECT_TRUE(CheckMessage("hello there", false, lists).empty()); +} From cef170ae47598905adb46229cb8b7ecbd0032f82 Mon Sep 17 00:00:00 2001 From: Aaron Kimbrell Date: Wed, 30 Sep 2026 08:05:00 -0500 Subject: [PATCH 3/4] feat(dashboard): chat filter page shows phrases and the block list's state Blocked entries marked Phrase, no Allow for a phrase, phrase matches named in the message test, and the Word files section says when blocklist.dcf is in the old format and how to rebuild it from blocklist.txt. issue 215 Co-Authored-By: Claude Opus 5.5 --- dDashboardServer/static/js/chat-filter.js | 27 ++++++++++++------- dDashboardServer/templates/chat_filter.jinja2 | 10 +++---- 2 files changed, 23 insertions(+), 14 deletions(-) diff --git a/dDashboardServer/static/js/chat-filter.js b/dDashboardServer/static/js/chat-filter.js index b85226f2c..5852fb287 100644 --- a/dDashboardServer/static/js/chat-filter.js +++ b/dDashboardServer/static/js/chat-filter.js @@ -8,6 +8,9 @@ var page = document.getElementById('chatFilterPage'); var canChat = !!page.dataset.canChat; var PAGE = 50; + function isPhrase(word) { return String(word).indexOf(' ') !== -1; } + function phraseBadge(word) { return isPhrase(word) ? ' ' + fmt.badge('Phrase', 'secondary') : ''; } + var REASONS = { blocked_here: 'Blocked on the staff list', allowed_here: 'Allowed on the staff list', @@ -26,7 +29,8 @@ function wordAction(w) { if (!w.word) return ''; - if (w.reason === 'blocked_here' || w.reason === 'allowed_here') return ''; + if (w.reason === 'blocked_here' || w.reason === 'allowed_here') return ''; + if (w.phrase) return ''; if (w.reason === 'not_allowed') return ''; return ''; } @@ -48,7 +52,7 @@ }).join(''); document.getElementById('testRows').innerHTML = d.words.map(function (w) { return '' + esc(w.word || '(empty: two spaces in a row)') + '' + (w.stopped ? fmt.badge('Stopped', 'danger') : fmt.badge('OK', 'success')) + '' + - '' + esc(REASONS[w.reason] || w.reason) + '' + wordAction(w) + ''; + '' + esc(REASONS[w.reason] || w.reason) + (w.phrase ? ' (the phrase ' + esc(w.phrase) + ')' : '') + '' + wordAction(w) + ''; }).join(''); document.getElementById('testResult').classList.remove('d-none'); }).catch(function () {}); @@ -83,10 +87,10 @@ var shown = rows.slice(listStart, listStart + PAGE); document.getElementById('listRows').innerHTML = shown.map(function (w) { var other = w.allowed ? 'blocked' : 'allowed'; - return '' + esc(w.word) + '' + (w.allowed ? fmt.badge('Allowed', 'success') : fmt.badge('Blocked', 'danger')) + '' + + return '' + esc(w.word) + phraseBadge(w.word) + '' + (w.allowed ? fmt.badge('Allowed', 'success') : fmt.badge('Blocked', 'danger')) + '' + '' + esc(w.added_by) + '' + esc(fmt.unix(w.added_at)) + '' + ' ' + - ' ' + + (isPhrase(w.word) && !w.allowed ? '' : ' ') + ''; }).join('') || '' + (words.length ? 'No words match.' : 'No words added yet.') + ''; document.getElementById('listRange').textContent = rows.length ? (listStart + 1) + '–' + (listStart + shown.length) + ' of ' + rows.length : ''; @@ -114,8 +118,9 @@ addForm.addEventListener('click', function (e) { var b = e.target.closest('button[data-list]'); if (b) addList = b.dataset.list; }); addForm.addEventListener('submit', function (e) { e.preventDefault(); - var word = addWord.value.trim(); + var word = addWord.value.trim().replace(/\s+/g, ' '); if (!word) return; + if (addList === 'allowed' && isPhrase(word)) { toast('Phrases can only be blocked: normal chat checks each word on its own, so allow the words instead', 'warning'); return; } confirmAdd(word, addList === 'allowed').then(function (done) { if (done) addWord.value = ''; }); }); @@ -129,7 +134,7 @@ var parts = [d.inAllowFile ? 'in chatplus_en_us.txt' : 'not in chatplus_en_us.txt']; if (d.inBlockFile) parts.push('in blocklist.dcf'); parts.push(d.dashboard ? (d.dashboard === 'blocked' ? 'on the Blocked list' : 'on the Allowed list') : 'on neither staff list'); - return '' + esc(d.word) + ' now: ' + esc(parts.join(', ')) + '.'; + return '' + esc(d.word) + '' + phraseBadge(d.word) + ' now: ' + esc(parts.join(', ')) + '.'; } function openConfirm(options) { @@ -186,7 +191,8 @@ function confirmAdd(word, allowed) { var p = openConfirm({ title: (allowed ? 'Allow "' : 'Block "') + word + '"?', - text: allowed ? 'Players may use it in normal chat. Running worlds apply it at once.' : 'It is stopped in all chat, even where a word file allows it. Running worlds apply it at once.', + text: allowed ? 'Players may use it in normal chat. Running worlds apply it at once.' + : (isPhrase(word) ? 'The phrase is stopped in all chat when its words come in a row. Running worlds apply it at once.' : 'It is stopped in all chat, even where a word file allows it. Running worlds apply it at once.'), tone: allowed ? 'success' : 'danger', button: allowed ? 'Allow' : 'Block', run: function () { return api.action('/api/chat_filter/words', { word: word, allowed: allowed }).then(function (d) { toast(d.message, 'success'); }); } @@ -228,8 +234,11 @@ ? esc(d.allowTotal) + ' words normal chat may use (from the client); ' + esc(d.onDashboard) + ' are on a staff list too, which wins.' : 'could not be read. Set client_location.') + '' + '
  • ' + esc(d.blockFile) + ': ' + (d.blockFileFound - ? esc(d.blockTotal) + ' words best friends\' free chat may not use. Stored as hashes, so they can\'t be listed; test a word to see if it is one.' - : 'not found next to the servers, so best friends\' free chat stops everything.') + '
  • '; + ? esc(d.blockTotal) + ' words' + (d.blockMaxWords > 1 ? ' and phrases (up to ' + esc(d.blockMaxWords) + ' words)' : '') + ' best friends\' free chat may not use. Stored as hashes, so they can\'t be listed; test a word or phrase to see if it is one.' + : d.blockFileOld + ? 'in the old format (its hashes depended on the platform), so the servers can\'t read it and best friends\' free chat stops everything.' + : 'not found next to the servers (' + esc(d.blockFileStatus) + '), so best friends\' free chat stops everything.') + + ' To change it, put the words in ' + esc(d.blockText) + ' next to the servers (one word or phrase per line) and start the servers again; they build ' + esc(d.blockFile) + ' from it.'; document.getElementById('fileWords').innerHTML = d.words.map(function (w) { var cls = w.dashboard === 'blocked' ? 'text-bg-danger' : w.dashboard === 'allowed' ? 'text-bg-success' : 'text-bg-secondary'; return ''; diff --git a/dDashboardServer/templates/chat_filter.jinja2 b/dDashboardServer/templates/chat_filter.jinja2 index c281b27ea..7caae932d 100644 --- a/dDashboardServer/templates/chat_filter.jinja2 +++ b/dDashboardServer/templates/chat_filter.jinja2 @@ -54,12 +54,12 @@
    -
    Blocked
    Stopped in all chat, even if a word file allows it.
    -
    Allowed
    Usable in normal chat, like the words in chatplus_en_us.txt.
    +
    Blocked
    Stopped in all chat, even if a word file allows it. A phrase (several words) is stopped when its words come in a row, whatever the spaces and punctuation between them.
    +
    Allowed
    Usable in normal chat, like the words in chatplus_en_us.txt. Single words only: normal chat checks each word on its own, as the client does.
    -
    -
    +
    +
    @@ -74,7 +74,7 @@
    - +
    WordListAdded byAdded
    Word or phraseListAdded byAdded
    From 9658014b4f71d978efae7af019184d018157c94c Mon Sep 17 00:00:00 2001 From: Aaron Kimbrell Date: Wed, 30 Sep 2026 08:05:26 -0500 Subject: [PATCH 4/4] docs(chat-filter): block list file, portable format, phrases How to build blocklist.dcf from blocklist.txt and where it goes, the version 3 layout and hash, what happens to old files, and blocked phrases; README "This branch" line. issue 215 Co-Authored-By: Claude Opus 5.5 --- README.md | 6 +++++- docs/Dashboard.md | 36 ++++++++++++++++++++++++++++++------ docs/IssueTracker.md | 1 + 3 files changed, 36 insertions(+), 7 deletions(-) diff --git a/README.md b/README.md index 28728013d..8fce7d677 100644 --- a/README.md +++ b/README.md @@ -137,6 +137,10 @@ locally. See [docs/UgcServer.md](docs/UgcServer.md). * **World hot reload:** worlds report the zone files they loaded (`.luz`, `.lvl`, triggers, terrain, navmesh); when one changes on disk, or on `/reloadworld` or the dashboard's Reload, master replaces those instances with new ones and moves their players over; properties are kept until empty instead ([docs/WorldHotReload.md](docs/WorldHotReload.md)). +* **Chat filter:** the block list and the allowed words cache (`.dcf`) are hashed with 64-bit FNV-1a, so a list works + on every platform (before, `std::hash` values made on one system never matched on another); old files are refused + with a log line. Servers build `blocklist.dcf` from a plain `blocklist.txt` next to them, and blocked entries can be + phrases. See "Block list file" in [docs/Dashboard.md](docs/Dashboard.md). * The chat server's old web API is removed; the dashboard's API covers online players, teams and announcements. ## License @@ -375,7 +379,7 @@ All listed files are required for a server to start. * masterconfig.ini * WorldServer(.exe) * worldconfig.ini -* blocklist.dcf +* blocklist.dcf (or blocklist.txt, one blocked word or phrase per line, which the servers build it from) * migrations * vanity * navmeshes diff --git a/docs/Dashboard.md b/docs/Dashboard.md index ce2934753..cc3f42889 100644 --- a/docs/Dashboard.md +++ b/docs/Dashboard.md @@ -1213,18 +1213,42 @@ character's owner sees only what is still in their mailbox. Deleting a character The **Chat Filter** page (GM 5+, `chat_filter_manage`, under Moderation) decides which words players below GM 2 may use in chat. The filter's files: `chatplus_en_us.txt` (client `res` folder) lists the words normal chat may use, -`blocklist.dcf` (next to the servers, hashes only) the words best friends' free chat may not. Approved character names -also count as allowed. Words are compared lower case, without `! ? ; . ,`. Changes apply at once in running worlds and in -the chat server's web chat; servers that start later read them. Changes are audited and go to the `moderation` webhook -event. +`blocklist.dcf` (next to the servers, hashes only) the words and phrases best friends' free chat may not. Approved +character names also count as allowed. Words are compared lower case, without `! ? ; . ,`. Changes apply at once in +running worlds and in the chat server's web chat; servers that start later read them. Changes are audited and go to the +`moderation` webhook event. + +Blocked entries can be phrases: a phrase is stopped when its words come in a row in a message, whatever the spaces and +punctuation between them, and the whole phrase is marked. Allowed entries are single words only, because normal +(whitelist) chat checks each word on its own, as the client does. + +#### Block list file + +`blocklist.dcf` is DLU's own file (the client reads no `.dcf` and doesn't hash chat words). To make or change it, put the +blocked words in `blocklist.txt` next to the servers (the build folder, beside `blocklist.dcf`): one word or phrase per +line, any case, punctuation `! ? ; . ,` ignored, blank lines skipped. The world and chat servers rebuild `blocklist.dcf` +from it when they start and it is newer than the `.dcf` (or the `.dcf` is missing or unreadable), then log how many +entries it has. With `dont_generate_dcf=1` they read `blocklist.txt` directly and write no file. `blocklist.txt` can be +removed afterwards; only the `.dcf` is needed. + +Format (little-endian): `uint32` magic `DCFB`, `uint32` version `3`, `uint32` most words in one entry, `uint64` count, +then that many `uint64` hashes, sorted. Each hash is 64-bit FNV-1a (offset basis `0xcbf29ce484222325`, prime +`0x100000001b3`) over the entry's bytes: the words lower case (ASCII), without `! ? ; . ,`, joined by one space. The +same words give the same file on every platform. The allowed words cache, `chatplus_en_us.dcf` in the client's `res` +folder, uses the same format and is built from `chatplus_en_us.txt`. + +Version 2 files (older DLU) stored `std::hash` values, which differ between compilers and platforms, so a list made on +one system never matched on another (issue 215). Servers refuse them: an old `chatplus_en_us.dcf` is rebuilt from +the `.txt`, and an old `blocklist.dcf` is logged as unreadable (free chat then stops every message) until it is rebuilt +from `blocklist.txt`. The **Word files** section shows which it is. The page has three sections: - **Test a message**: type a message, pick normal or best friends' free chat, and see whether it would be sent and why, word by word (in the file, allowed or blocked here, a character name, not allowed, in the blocked words file). Each word has a Block, Allow or Remove button. -- **Staff lists**: **Blocked** words are stopped in all chat, even where a file allows them; **Allowed** words are usable - in normal chat. Search, filter by list, 50 per page. Block, Allow (or move to the other list) and Remove each open a +- **Staff lists**: **Blocked** words and phrases are stopped in all chat, even where a file allows them (phrases are + marked **Phrase**); **Allowed** words are usable in normal chat. Search, filter by list, 50 per page. Block, Allow (or move to the other list) and Remove each open a confirmation that shows where the word stands now and, for Block and Allow, the recent chat it changes (players' chat containing it that would have been stopped, or stopped messages containing it; the newest 1000 messages with the text; needs `chat_view`). diff --git a/docs/IssueTracker.md b/docs/IssueTracker.md index 48ea6c05c..24167024a 100644 --- a/docs/IssueTracker.md +++ b/docs/IssueTracker.md @@ -10,6 +10,7 @@ State: **done** = fixed on this branch, needs an in-game check; **partial** = pa |---|---|---|---| | 159 | BUG: Brick-by-brick models are deleted instead of put away | done | `6ce261c5` fix: brick by brick and model placement work the way the client expects | | 185 | BUG: Assembly Engineer Fortress Knockback | partial | `2d76c81b` feat: server side knockback for AI moved objects | +| 215 | ENH: Bring chat filter closer to Live | partial | `c1bcda8d` fix(chat-filter): portable .dcf hashing, block list phrases; `b279d187` shipped blocklist.dcf in the portable format; `cef170ae` dashboard phrases. The block list works on every platform and takes phrases; the rest of the issue is open | | 225 | ENH: "bind_ip" config option | done | `e36f894f` feat: bind_ip setting for the server sockets | | 257 | EH: Crux Prime shields stun instead of knockback | done | `2d76c81b` feat: server side knockback for AI moved objects | | 307 | Spider Queen scream on spiderling death | done | `4900f11c` fix(scripts): the Spider Queen screams from the mountain when a spiderling dies (issue 307) |