|
SciLex
A header-only C++20 lexer built on REAL
|
A token rule: a kind, the pattern that recognizes it, whether matches are discarded (whitespace, comments), and — for contextual lexing — the modes it is active in and an optional mode transition it fires when it wins. More...
#include <lexer.hpp>
Public Attributes | |
| int | kind |
| Kind assigned to tokens this rule produces. | |
| real::regex | pattern |
| The recognizer (a linear-time REAL regex; its flags are the author's — see above). | |
| bool | skip {false} |
| If true, matches are consumed but not emitted. | |
| std::vector< std::string > | in_mode {} |
| Modes this rule is active in; empty ⇒ {"default"}. | |
| std::optional< mode_action > | action {} |
| Mode transition fired when this rule wins. | |
A token rule: a kind, the pattern that recognizes it, whether matches are discarded (whitespace, comments), and — for contextual lexing — the modes it is active in and an optional mode transition it fires when it wins.
in_mode empty means the rule is active in the implicit "default" mode only, so a plain {kind, pattern, skip} rule keeps working unchanged.
The pattern is a fully-formed real::regex, so the grammar author owns its flags.
This is a real trade-off, not a footnote. \w+ (or [^\W\d]\w*) with the default flags reads Unicode identifiers — café, 変数 — the faithful behaviour for a language like Python 3. But a Unicode \w expands into more UTF-8 byte transitions than a DFA is built from, and \b is a zero-width assertion no DFA represents, so a rule holding either stays on the general Pike engine while the mode's other rules take the DFA (same tokens). That rule is tried at every position its first byte allows, so it keeps a share of the per-rule cost. The narrower Unicode \d and \s expand and stay on the DFA. BENCHMARKS.md measures the cost per grammar: the python-unicode grammar, whose identifier rule stays on Pike, keeps a smaller part of the DFA's gain than python, whose rules all take the DFA — the Unicode identifier costs part of the fast path.
If your identifiers are ASCII by specification (JSON, SQL, C), pin (?a) inline in the pattern (or pass real::flags::ascii) to keep \w \d \s \b ASCII and small, DFA-representable, and fast — this is what the examples/ grammars do. If you want Unicode identifiers, write \w+ and accept the general-engine floor. The two spellings tokenize the same ASCII input identically; they differ only on non-ASCII input and on whether the mode can be a DFA.
| std::optional<mode_action> scilex::rule::action {} |
Mode transition fired when this rule wins.
| std::vector<std::string> scilex::rule::in_mode {} |
Modes this rule is active in; empty ⇒ {"default"}.
| int scilex::rule::kind |
Kind assigned to tokens this rule produces.
| real::regex scilex::rule::pattern |
The recognizer (a linear-time REAL regex; its flags are the author's — see above).
| bool scilex::rule::skip {false} |
If true, matches are consumed but not emitted.