SciLex
A header-only C++20 lexer built on REAL
Loading...
Searching...
No Matches
Classes | Public Member Functions | Friends | List of all members
scilex::lexer Class Reference

A lexer built from an ordered list of rules. More...

#include <lexer.hpp>

Public Member Functions

 lexer (std::vector< rule > rules, std::vector< std::string > insignificant_modes={}, std::vector< std::string > dfa_modes={}, error_policy errors=error_policy::raise, column_unit columns=column_unit::bytes, dfa_policy dfa=dfa_policy::automatic)
 Builds a lexer from rules (taken by value, then moved in).
 
column_unit columns () const noexcept
 The unit this lexer counts position::column in (positions do not carry it, so a consumer that needs to interpret a column reads the unit here).
 
std::vector< token > tokenize (std::string_view source, eof_policy policy=eof_policy::omit) const
 Tokenizes source into the sequence of non-skipped tokens.
 
token_range scan (std::string_view source, eof_policy policy=eof_policy::omit) const &
 Returns a lazy range over the non-skipped tokens of source.
 
token_range scan (std::string_view source, eof_policy policy=eof_policy::omit) const &&=delete
 Deleted: the range would point into a temporary lexer.
 
token_stream stream () const &
 A stream: text fed in pieces, the tokens returned as soon as no text still to come can change them, and exactly the tokens tokenize gives the whole text (see token_stream).
 
token_stream stream () const &&=delete
 Deleted: the stream would point into a temporary lexer.
 
const std::vector< bool > & mode_significant () const noexcept
 The per-mode-id layout-significance policy (see scilex::layout). Index by a token's mode_id; false marks an insignificant mode. Empty unless the lexer was built with insignificant_modes.
 
const std::string & mode_name (std::size_t id) const noexcept
 The name of mode id (0 is "default"), for labelling tokens.
 
std::vector< std::string > dfa_modes_active () const
 The modes actually accelerated by a DFA fast path.
 
std::vector< std::size_t > pike_rules (const std::string &mode) const
 The rules of mode that run on the per-rule Pike path, by index into the rules the lexer was built from: every rule of a mode with no DFA, and in an accelerated mode the rules its DFA cannot take (see dfa_modes_active).
 
position end_of (std::string_view source, const token &tok) const
 Where tok ends in source: the position just past its last byte, with the line and column this lexer's scilex::column_unit gives it — exactly where the scan's cursor stood after the token.
 

Friends

class token_iterator
 
class token_stream
 

Detailed Description

A lexer built from an ordered list of rules.

Order matters only as a tie-breaker between rules whose matches have equal length (the first such rule wins). Put more specific rules (keywords) before their general counterparts (identifiers).

Thread safety
A lexer is immutable once built: its DFAs are built in the constructor, and every tokenize call and every scan range keeps its own mode stack and walk memos. One const lexer may therefore be shared by any number of threads, each tokenizing its own source. A token_iterator and a token_stream are cursors: drive each one from a single thread. Rules left on Pike (pike_rules) call real::regex, whose lazy DFAs each thread leases from the regex's own pool (REAL 2026.9.8 and later), so threads scanning with such a rule take no lock.

Constructor & Destructor Documentation

◆ lexer()

scilex::lexer::lexer ( std::vector< rule >  rules,
std::vector< std::string >  insignificant_modes = {},
std::vector< std::string >  dfa_modes = {},
error_policy  errors = error_policy::raise,
column_unit  columns = column_unit::bytes,
dfa_policy  dfa = dfa_policy::automatic 
)
inlineexplicit

Builds a lexer from rules (taken by value, then moved in).

Parameters
[in]rulesThe ordered token rules.
[in]insignificant_modesModes whose tokens carry no layout structure (Layout Awareness Level A — see scilex::layout). Each name must be a mode the rules use; empty (the default) leaves every mode significant, so mode_significant has no effect.
[in]dfa_modesModes to accelerate with a real::dfa fast path (one DFA pass replaces the per-rule Pike dispatch) under dfa_policy::requested; under the default dfa_policy::automatic every mode is tried and this list adds nothing. Each name must be a mode the rules use. The attempt is best-effort: a rule that cannot be a DFA (a zero-width assertion) or whose DFA would change an answer (its match() is not its longest match, such as as|assert or a lazy delimiter) silently stays on Pike beside the mode's DFA — see dfa_modes_active. The token stream is identical either way: the decision is exact, not sampled.
[in]errorsWhat to do at a byte no rule can lex: error_policy::raise (the default — throw) or error_policy::token (recover, emitting an scilex::error token). The recovery path never throws per byte; the token stream under raise is unchanged.
[in]columnsThe unit each token's position::column is counted in: column_unit::bytes (the default, unchanged), column_unit::codepoints, or column_unit::utf16. The unit is not stored on the position — read it back with columns().
[in]dfaWhich modes are tried for DFA acceleration (dfa_policy); the default tries all of them.
Exceptions
std::invalid_argumentIf a rule's kind is reserved (end_of_input to scilex::error), a transition rule is malformed (empty pattern or target), or insignificant_modes / dfa_modes names an unknown mode.

Member Function Documentation

◆ columns()

column_unit scilex::lexer::columns ( ) const
inlinenoexcept

The unit this lexer counts position::column in (positions do not carry it, so a consumer that needs to interpret a column reads the unit here).

◆ dfa_modes_active()

std::vector< std::string > scilex::lexer::dfa_modes_active ( ) const
inline

The modes actually accelerated by a DFA fast path.

A mode is listed when at least one of its rules runs on the DFA. A rule the DFA cannot take — it needs an assertion no DFA represents (real::dfa_error), or its match() is not its longest match — stays on Pike beside the DFA; a mode where no rule can go on one is absent here. The tokens are the same either way: acceleration is an optimizer, not a guarantee.

◆ end_of()

position scilex::lexer::end_of ( std::string_view  source,
const token &  tok 
) const
inline

Where tok ends in source: the position just past its last byte, with the line and column this lexer's scilex::column_unit gives it — exactly where the scan's cursor stood after the token.

A token carries its start and its lexeme, not its end, so the token stream costs no more for the callers that never ask. A zero-width token (a synthetic newline, indent or dedent from layout, the end-of-input token) ends where it starts.

Parameters
[in]sourceThe text tok was lexed from; its lexeme must view into it.
[in]tokA token this lexer produced from source.
Returns
The position after tok.
Exceptions
std::invalid_argumentIf tok's lexeme is not the bytes of source at its start offset.

◆ mode_name()

const std::string & scilex::lexer::mode_name ( std::size_t  id) const
inlinenoexcept

The name of mode id (0 is "default"), for labelling tokens.

◆ mode_significant()

const std::vector< bool > & scilex::lexer::mode_significant ( ) const
inlinenoexcept

The per-mode-id layout-significance policy (see scilex::layout). Index by a token's mode_id; false marks an insignificant mode. Empty unless the lexer was built with insignificant_modes.

◆ pike_rules()

std::vector< std::size_t > scilex::lexer::pike_rules ( const std::string &  mode) const
inline

The rules of mode that run on the per-rule Pike path, by index into the rules the lexer was built from: every rule of a mode with no DFA, and in an accelerated mode the rules its DFA cannot take (see dfa_modes_active).

Parameters
[in]modeA mode the rules use.
Returns
The indices, ascending.
Exceptions
std::invalid_argumentIf mode is not a mode the rules use.

◆ scan() [1/2]

token_range scilex::lexer::scan ( std::string_view  source,
eof_policy  policy = eof_policy::omit 
) const &
inline

Returns a lazy range over the non-skipped tokens of source.

Each ++ produces the next token on demand; nothing but the current token is held. Usable in a range-for. Errors surface as lex_error thrown while advancing.

Parameters
[in]sourceThe text to scan (must outlive the iteration; each token's lexeme views into it).
[in]policyWhether to yield a terminal end_of_input token.
Returns
A token_range whose iterators yield token values.
Exceptions
lex_error(while iterating) if some position matches no rule.

◆ scan() [2/2]

token_range scilex::lexer::scan ( std::string_view  source,
eof_policy  policy = eof_policy::omit 
) const &&
delete

Deleted: the range would point into a temporary lexer.

◆ stream() [1/2]

token_stream scilex::lexer::stream ( ) const &
inline

A stream: text fed in pieces, the tokens returned as soon as no text still to come can change them, and exactly the tokens tokenize gives the whole text (see token_stream).

Returns
A stream over this lexer, which must outlive it.

◆ stream() [2/2]

token_stream scilex::lexer::stream ( ) const &&
delete

Deleted: the stream would point into a temporary lexer.

◆ tokenize()

std::vector< token > scilex::lexer::tokenize ( std::string_view  source,
eof_policy  policy = eof_policy::omit 
) const
inline

Tokenizes source into the sequence of non-skipped tokens.

At each position every rule is matched anchored; the longest match wins (ties broken by rule order). A zero-length winning match (a nullable rule with no longer match here) cannot advance the scan, so it is reported as a lex_error rather than allowed to stall.

Parameters
[in]sourceThe text to tokenize (must outlive the returned tokens; each token's lexeme views into it).
[in]policyWhether to append a terminal end_of_input token.
Returns
The tokens in source order, skip-rule matches omitted.
Exceptions
lex_errorIf some position is matched by no rule, or only by a zero-length match.

Friends And Related Symbol Documentation

◆ token_iterator

friend class token_iterator
friend

◆ token_stream

friend class token_stream
friend

The documentation for this class was generated from the following file: