|
SciLex
A header-only C++20 lexer built on REAL
|
A lexer built from an ordered list of rules. More...
#include <lexer.hpp>
Public Member Functions | |
| lexer (std::vector< rule > rules, std::vector< std::string > insignificant_modes={}, std::vector< std::string > dfa_modes={}, error_policy errors=error_policy::raise, column_unit columns=column_unit::bytes, dfa_policy dfa=dfa_policy::automatic) | |
Builds a lexer from rules (taken by value, then moved in). | |
| column_unit | columns () const noexcept |
The unit this lexer counts position::column in (positions do not carry it, so a consumer that needs to interpret a column reads the unit here). | |
| std::vector< token > | tokenize (std::string_view source, eof_policy policy=eof_policy::omit) const |
Tokenizes source into the sequence of non-skipped tokens. | |
| token_range | scan (std::string_view source, eof_policy policy=eof_policy::omit) const & |
Returns a lazy range over the non-skipped tokens of source. | |
| token_range | scan (std::string_view source, eof_policy policy=eof_policy::omit) const &&=delete |
| Deleted: the range would point into a temporary lexer. | |
| token_stream | stream () const & |
| A stream: text fed in pieces, the tokens returned as soon as no text still to come can change them, and exactly the tokens tokenize gives the whole text (see token_stream). | |
| token_stream | stream () const &&=delete |
| Deleted: the stream would point into a temporary lexer. | |
| const std::vector< bool > & | mode_significant () const noexcept |
The per-mode-id layout-significance policy (see scilex::layout). Index by a token's mode_id; false marks an insignificant mode. Empty unless the lexer was built with insignificant_modes. | |
| const std::string & | mode_name (std::size_t id) const noexcept |
The name of mode id (0 is "default"), for labelling tokens. | |
| std::vector< std::string > | dfa_modes_active () const |
| The modes actually accelerated by a DFA fast path. | |
| std::vector< std::size_t > | pike_rules (const std::string &mode) const |
The rules of mode that run on the per-rule Pike path, by index into the rules the lexer was built from: every rule of a mode with no DFA, and in an accelerated mode the rules its DFA cannot take (see dfa_modes_active). | |
| position | end_of (std::string_view source, const token &tok) const |
Where tok ends in source: the position just past its last byte, with the line and column this lexer's scilex::column_unit gives it — exactly where the scan's cursor stood after the token. | |
Friends | |
| class | token_iterator |
| class | token_stream |
A lexer built from an ordered list of rules.
Order matters only as a tie-breaker between rules whose matches have equal length (the first such rule wins). Put more specific rules (keywords) before their general counterparts (identifiers).
const lexer may therefore be shared by any number of threads, each tokenizing its own source. A token_iterator and a token_stream are cursors: drive each one from a single thread. Rules left on Pike (pike_rules) call real::regex, whose lazy DFAs each thread leases from the regex's own pool (REAL 2026.9.8 and later), so threads scanning with such a rule take no lock.
|
inlineexplicit |
Builds a lexer from rules (taken by value, then moved in).
| [in] | rules | The ordered token rules. |
| [in] | insignificant_modes | Modes whose tokens carry no layout structure (Layout Awareness Level A — see scilex::layout). Each name must be a mode the rules use; empty (the default) leaves every mode significant, so mode_significant has no effect. |
| [in] | dfa_modes | Modes to accelerate with a real::dfa fast path (one DFA pass replaces the per-rule Pike dispatch) under dfa_policy::requested; under the default dfa_policy::automatic every mode is tried and this list adds nothing. Each name must be a mode the rules use. The attempt is best-effort: a rule that cannot be a DFA (a zero-width assertion) or whose DFA would change an answer (its match() is not its longest match, such as as|assert or a lazy delimiter) silently stays on Pike beside the mode's DFA — see dfa_modes_active. The token stream is identical either way: the decision is exact, not sampled. |
| [in] | errors | What to do at a byte no rule can lex: error_policy::raise (the default — throw) or error_policy::token (recover, emitting an scilex::error token). The recovery path never throws per byte; the token stream under raise is unchanged. |
| [in] | columns | The unit each token's position::column is counted in: column_unit::bytes (the default, unchanged), column_unit::codepoints, or column_unit::utf16. The unit is not stored on the position — read it back with columns(). |
| [in] | dfa | Which modes are tried for DFA acceleration (dfa_policy); the default tries all of them. |
| std::invalid_argument | If a rule's kind is reserved (end_of_input to scilex::error), a transition rule is malformed (empty pattern or target), or insignificant_modes / dfa_modes names an unknown mode. |
|
inlinenoexcept |
The unit this lexer counts position::column in (positions do not carry it, so a consumer that needs to interpret a column reads the unit here).
|
inline |
The modes actually accelerated by a DFA fast path.
A mode is listed when at least one of its rules runs on the DFA. A rule the DFA cannot take — it needs an assertion no DFA represents (real::dfa_error), or its match() is not its longest match — stays on Pike beside the DFA; a mode where no rule can go on one is absent here. The tokens are the same either way: acceleration is an optimizer, not a guarantee.
Where tok ends in source: the position just past its last byte, with the line and column this lexer's scilex::column_unit gives it — exactly where the scan's cursor stood after the token.
A token carries its start and its lexeme, not its end, so the token stream costs no more for the callers that never ask. A zero-width token (a synthetic newline, indent or dedent from layout, the end-of-input token) ends where it starts.
| [in] | source | The text tok was lexed from; its lexeme must view into it. |
| [in] | tok | A token this lexer produced from source. |
tok. | std::invalid_argument | If tok's lexeme is not the bytes of source at its start offset. |
|
inlinenoexcept |
The name of mode id (0 is "default"), for labelling tokens.
|
inlinenoexcept |
The per-mode-id layout-significance policy (see scilex::layout). Index by a token's mode_id; false marks an insignificant mode. Empty unless the lexer was built with insignificant_modes.
|
inline |
The rules of mode that run on the per-rule Pike path, by index into the rules the lexer was built from: every rule of a mode with no DFA, and in an accelerated mode the rules its DFA cannot take (see dfa_modes_active).
| [in] | mode | A mode the rules use. |
| std::invalid_argument | If mode is not a mode the rules use. |
|
inline |
Returns a lazy range over the non-skipped tokens of source.
Each ++ produces the next token on demand; nothing but the current token is held. Usable in a range-for. Errors surface as lex_error thrown while advancing.
| [in] | source | The text to scan (must outlive the iteration; each token's lexeme views into it). |
| [in] | policy | Whether to yield a terminal end_of_input token. |
| lex_error | (while iterating) if some position matches no rule. |
|
delete |
Deleted: the range would point into a temporary lexer.
|
inline |
A stream: text fed in pieces, the tokens returned as soon as no text still to come can change them, and exactly the tokens tokenize gives the whole text (see token_stream).
|
delete |
Deleted: the stream would point into a temporary lexer.
|
inline |
Tokenizes source into the sequence of non-skipped tokens.
At each position every rule is matched anchored; the longest match wins (ties broken by rule order). A zero-length winning match (a nullable rule with no longer match here) cannot advance the scan, so it is reported as a lex_error rather than allowed to stall.
| [in] | source | The text to tokenize (must outlive the returned tokens; each token's lexeme views into it). |
| [in] | policy | Whether to append a terminal end_of_input token. |
| lex_error | If some position is matched by no rule, or only by a zero-length match. |
|
friend |
|
friend |