# Pattern Matching Operations Text pattern matching and searching operations with support for different UTF encodings. The definition and the string to match can use different UTF encodings. `Regex_` accepts `ObjectString` as the definition. `MatchHead`, `Match`, `TestHead`, `Test`, `Search`, `Split` and `Cut` accept `ObjectString` or `const U*` text to match with the regular expression. `Search`, `Split` and `Cut` append results to a caller-provided `RegexMatch_::List&`; they do not return a collection. `Split` and `Cut` also take `keepEmptyMatch` to decide whether empty unmatched fragments are included. ## Core Pattern Matching Methods ### MatchHead `MatchHead` matches a prefix of the string. Pure mode chooses the longest match; rich mode follows the pattern's branch and repetition priorities (`Source/Regex/RegexRich.cpp` and `Source/Regex/RegexPure.cpp`). This method attempts to match the pattern starting from the beginning of the input string. Lazy quantifiers in rich mode can return a shorter match. It will return detailed match information including captured groups if the pattern matches. ### Match `Match` finds the earliest substring which matches the regular expression. This method searches through the entire input string to find the first occurrence of a substring that matches the pattern. Unlike `MatchHead`, it doesn't require the match to be at the beginning of the string. ### TestHead `TestHead` performs a similar action to `MatchHead`, but it only returns `bool` without detailed information. This is an optimization when you only need to know whether the beginning of the string matches the pattern, without requiring access to the actual match data or captured groups. ### Test `Test` performs a similar action to `Match`, but it only returns `bool` without detailed information. This is an optimization when you only need to know whether the string contains a substring that matches the pattern, without requiring access to the actual match data or captured groups. ## Advanced Pattern Operations ### Search `Search` finds all substrings which match the regular expression. All results do not overlap with each other. This method appends all non-overlapping successful matches found in the input string to `RegexMatch_::List&`. When a match is found, searching continues after that match. ### Split `Split` uses the regular expression as a splitter, finding all remaining substrings. This method treats the pattern as a delimiter and appends the parts between successful matches. The appended `RegexMatch_` objects have `Success()` equal to `false`. `keepEmptyMatch` controls whether empty unmatched fragments are appended. ### Cut `Cut` combines both `Search` and `Split`, finding all substrings in order, regardless if one matches or not. This method appends all parts of the string in sequence, both the parts that match the pattern and the parts that do not match. `Success()` distinguishes successful pattern matches from unmatched fragments, and `keepEmptyMatch` controls empty unmatched fragments. ## Extra Content ### UTF Encoding Support One of the key features of VlppRegex is its support for different UTF encodings between the pattern definition and the input text: - The regex pattern is defined using `Regex_` where `T` is the character type for the pattern - The input text uses type `U` in the matching methods, through `ObjectString` or `const U*` - This allows patterns defined in one encoding to match text in another encoding - Supported character types include `wchar_t`, `char8_t`, `char16_t`, `char32_t` ### Performance Considerations The VlppRegex engine has specific performance characteristics: - **DFA Compatible vs Incompatible**: Backreferences, lookahead, anchors, and lazy loops require rich mode. Captures require rich mode for detailed matches, but otherwise compatible patterns can still use DFA mode for `Test` and `TestHead` (`Source/Regex/Regex.cpp`). - **Escaping Optimization**: Using `/` instead of `\` for escaping can improve readability in C++ code - **Method Selection**: Choose simpler methods like `Test` or `TestHead` when you only need boolean results - **Mode Inspection**: Use `IsPureMatch()` and `IsPureTest()` to check whether DFA mode is used for matching and testing ### Syntax Differences from .NET While mostly compatible with .NET regex syntax, VlppRegex has important differences: - **Dot Character**: `.` matches literal '.' character, while `/.` or `\.` matches any character - **Escaping**: Both `/` and `\` perform escaping (prefer `/` for C++ compatibility) - **Character Classes**: `\s`, `\S`, `\d`, `\D`, `\l`, `\L`, `\w` and `\W` are supported, and each one can also be written with `/` - **Quantifiers**: Standard quantifiers (`*`, `+`, `?`, `{n,m}`) work as expected ### Error Handling When using regex operations: - Syntax errors throw `RegexException` during `Regex_` construction (`Source/Regex/AST/RegexParser.cpp`) - `MatchHead` and `Match` return `nullptr` when no match is found - `TestHead` and `Test` return `false` when no match is found - `Search` appends no items when no successful match is found ### Common Usage Patterns **Prefix test**: ```cpp Regex regex(L"/d+"); bool hasNumberPrefix = regex.TestHead(input); ``` **Extracting all matches**: ```cpp Regex regex(L"/w+"); RegexMatch::List matches; regex.Search(text, matches); for (auto match : matches) { // Process each word } ``` **Splitting text**: ```cpp Regex regex(L"[,;]"); RegexMatch::List parts; regex.Split(csvLine, false, parts); ``` **Complete decomposition**: ```cpp Regex regex(L"/d+"); RegexMatch::List parts; regex.Cut(mixedText, false, parts); // Appends both numbers and non-numbers ```