C++#

#include <scilex/scilex.hpp> brings in the lexer, rules and tokens; scilex/layout.hpp and scilex/grammar.hpp are opt-in.

The lexer#

class lexer#

A lexer built from an ordered list of rules.

Order matters only as a tie-breaker between rules whose matches have equal length (the first such rule wins). Put more specific rules (keywords) before their general counterparts (identifiers).

Thread safety

A lexer is immutable once built: its DFAs are built in the constructor, and every tokenize call and every scan range keeps its own mode stack and walk memos. One const lexer may therefore be shared by any number of threads, each tokenizing its own source. A token_iterator and a token_stream are cursors: drive each one from a single thread. Rules left on Pike (pike_rules) call real::regex, whose lazy DFAs each thread leases from the regex’s own pool (REAL 2026.9.8 and later), so threads scanning with such a rule take no lock.

Public Functions

inline explicit lexer(std::vector<rule> rules, std::vector<std::string> insignificant_modes = {}, std::vector<std::string> dfa_modes = {}, error_policy errors = error_policy::raise, column_unit columns = column_unit::bytes, dfa_policy dfa = dfa_policy::automatic)#

Builds a lexer from rules (taken by value, then moved in).

Parameters:
  • rules – [in] The ordered token rules.

  • insignificant_modes – [in] Modes whose tokens carry no layout structure (Layout Awareness Level A — see scilex::layout). Each name must be a mode the rules use; empty (the default) leaves every mode significant, so mode_significant has no effect.

  • dfa_modes – [in] Modes to accelerate with a real::dfa fast path (one DFA pass replaces the per-rule Pike dispatch) under dfa_policy::requested; under the default dfa_policy::automatic every mode is tried and this list adds nothing. Each name must be a mode the rules use. The attempt is best-effort: a rule that cannot be a DFA (a zero-width assertion) or whose DFA would change an answer (its match() is not its longest match, such as as|assert or a lazy delimiter) silently stays on Pike beside the mode’s DFA — see dfa_modes_active. The token stream is identical either way: the decision is exact, not sampled.

  • errors – [in] What to do at a byte no rule can lex: error_policy::raise (the default — throw) or error_policy::token (recover, emitting an scilex::error token). The recovery path never throws per byte; the token stream under raise is unchanged.

  • columns – [in] The unit each token’s position::column is counted in: column_unit::bytes (the default, unchanged), column_unit::codepoints, or column_unit::utf16. The unit is not stored on the position — read it back with columns().

  • dfa – [in] Which modes are tried for DFA acceleration (dfa_policy); the default tries all of them.

Throws:

std::invalid_argument – If a rule’s kind is reserved (end_of_input to scilex::error), a transition rule is malformed (empty pattern or target), or insignificant_modes / dfa_modes names an unknown mode.

inline column_unit columns() const noexcept#

The unit this lexer counts position::column in (positions do not carry it, so a consumer that needs to interpret a column reads the unit here).

inline std::vector<token> tokenize(std::string_view source, eof_policy policy = eof_policy::omit) const#

Tokenizes source into the sequence of non-skipped tokens.

At each position every rule is matched anchored; the longest match wins (ties broken by rule order). A zero-length winning match (a nullable rule with no longer match here) cannot advance the scan, so it is reported as a lex_error rather than allowed to stall.

Parameters:
  • source – [in] The text to tokenize (must outlive the returned tokens; each token’s lexeme views into it).

  • policy – [in] Whether to append a terminal end_of_input token.

Throws:

lex_error – If some position is matched by no rule, or only by a zero-length match.

Returns:

The tokens in source order, skip-rule matches omitted.

inline token_range scan(std::string_view source, eof_policy policy = eof_policy::omit) const &#

Returns a lazy range over the non-skipped tokens of source.

Each ++ produces the next token on demand; nothing but the current token is held. Usable in a range-for. Errors surface as lex_error thrown while advancing.

Parameters:
  • source – [in] The text to scan (must outlive the iteration; each token’s lexeme views into it).

  • policy – [in] Whether to yield a terminal end_of_input token.

Throws:

lex_error – (while iterating) if some position matches no rule.

Returns:

A token_range whose iterators yield token values.

token_range scan(std::string_view source, eof_policy policy = eof_policy::omit) const && = delete#

Deleted: the range would point into a temporary lexer.

inline const std::vector<bool> &mode_significant() const noexcept#

The per-mode-id layout-significance policy (see scilex::layout). Index by a token’s mode_id; false marks an insignificant mode. Empty unless the lexer was built with insignificant_modes.

inline const std::string &mode_name(std::size_t id) const noexcept#

The name of mode id (0 is “default”), for labelling tokens.

inline std::vector<std::string> dfa_modes_active() const#

The modes actually accelerated by a DFA fast path.

A mode is listed when at least one of its rules runs on the DFA. A rule the DFA cannot take — it needs an assertion no DFA represents (real::dfa_error), or its match() is not its longest match — stays on Pike beside the DFA; a mode where no rule can go on one is absent here. The tokens are the same either way: acceleration is an optimizer, not a guarantee.

inline std::vector<std::size_t> pike_rules(const std::string &mode) const#

The rules of mode that run on the per-rule Pike path, by index into the rules the lexer was built from: every rule of a mode with no DFA, and in an accelerated mode the rules its DFA cannot take (see dfa_modes_active).

Parameters:

mode – [in] A mode the rules use.

Throws:

std::invalid_argument – If mode is not a mode the rules use.

Returns:

The indices, ascending.

inline position end_of(std::string_view source, const token &tok) const#

Where tok ends in source: the position just past its last byte, with the line and column this lexer’s scilex::column_unit gives it — exactly where the scan’s cursor stood after the token.

A token carries its start and its lexeme, not its end, so the token stream costs no more for the callers that never ask. A zero-width token (a synthetic newline, indent or dedent from layout, the end-of-input token) ends where it starts.

Parameters:
  • source – [in] The text tok was lexed from; its lexeme must view into it.

  • tok – [in] A token this lexer produced from source.

Throws:

std::invalid_argument – If tok's lexeme is not the bytes of source at its start offset.

Returns:

The position after tok.

enum class scilex::eof_policy#

Whether tokenization appends a synthetic end-of-input token.

eof_policy::append yields one final token of kind end_of_input at the end position once the input is exhausted — the parser-friendly mode, so a cursor always has a current token to match against.

Values:

enumerator omit#

Stop at the last real token (default).

enumerator append#

Append one end_of_input token at the end position.

enum class scilex::error_policy#

What a lexer does when it reaches a byte that no rule in the active mode can begin.

The default preserves the historical behaviour exactly; token is opt-in recovery.

Values:

enumerator raise#

Throw a lex_error at the first unmatched byte (the default).

enumerator token#

Recover: emit the maximal unmatched byte run as one scilex::error token and resume. The cost of an error run is the grammar’s no-match cost: a first-byte pre-filter skips positions no rule can begin (usually O(1) per byte), so an unanchored, greedy rule that scans far before failing is what makes recovery expensive on a long run — prefer a definite leading byte.

enum class scilex::column_unit#

The unit a token’s position::column is counted in.

The default bytes is the historical behaviour, bit-for-bit. codepoints counts Unicode scalar values (each valid UTF-8 codepoint is one column), and utf16 counts UTF-16 code units (a BMP codepoint is 1, an astral codepoint 2) — the unit an LSP client expects. A malformed byte (an orphan continuation, an overlong or out-of-range sequence) counts as one unit in every mode, so the column stays defined on the error runs error_policy::token emits. The chosen unit is not carried on position (one field, not self-describing) — the lexer declares it via lexer::columns, a named trade-off rather than a silent default.

Values:

enumerator bytes#

One column per byte (the default; column == byte offset within the line + 1).

enumerator codepoints#

One column per Unicode scalar value (a valid UTF-8 codepoint).

enumerator utf16#

One column per UTF-16 code unit (BMP = 1, astral = 2) — the LSP unit.

enum class scilex::dfa_policy : std::uint8_t#

Which modes a lexer tries to accelerate with a real::dfa.

The token stream is the same either way: a mode keeps its DFA only where the lexer has decided that the DFA reproduces the per-rule munch on every input (see lexer::dfa_modes_active).

Values:

enumerator automatic#

Every mode (the default). In each, the rules a DFA reproduces go on one and the others stay on the per-rule path; the cost is the DFA construction.

enumerator requested#

Only the modes named in dfa_modes; an empty list keeps every mode on the per-rule path.

constexpr std::size_t scilex::max_mode_depth = {65536}#

The deepest the mode stack may grow. A push beyond it is a lex_error under either error policy: each frame costs memory, and without a bound an input of openers alone (16 MiB of ( under the python grammar) grew the stack past a gigabyte. 65 536 frames is ~2 MiB, far beyond any nesting a real grammar reaches.

Rules and modes#

struct rule#

A token rule: a kind, the pattern that recognizes it, whether matches are discarded (whitespace, comments), and — for contextual lexing — the modes it is active in and an optional mode transition it fires when it wins.

in_mode empty means the rule is active in the implicit “default” mode only, so a plain {kind, pattern, skip} rule keeps working unchanged.

The pattern is a fully-formed real::regex, so the grammar author owns its flags.

Public Members

int kind#

Kind assigned to tokens this rule produces.

real::regex pattern#

The recognizer (a linear-time REAL regex; its flags are the author’s — see above).

bool skip = {false}#

If true, matches are consumed but not emitted.

std::vector<std::string> in_mode = {}#

Modes this rule is active in; empty ⇒ {“default”}.

std::optional<mode_action> action = {}#

Mode transition fired when this rule wins.

struct mode_action#

A mode transition, fired when its rule wins, acting on the scan’s mode stack: enter a nested mode, leave the current one, or replace it.

Public Types

enum class op#

The kind of transition.

Values:

enumerator push#

Enter target, remembering the mode below it (a nested context).

enumerator pop#

Leave the current mode, returning to the one beneath it.

enumerator set#

Replace the current mode with target (stack depth unchanged).

Public Members

op operation#

Which transition to perform.

std::string target = {}#

The mode push/set enters; ignored (and omittable) for pop.

std::size_t target_id = {0}#

The interned id of target, resolved once when the lexer is built so the per-token transition is a field read, not a name→id map lookup. Internal cache: a caller leaves it at 0 and sets only target; pop leaves it unused.

Tokens and positions#

struct token#

One lexical token: a typed slice of the source.

Public Members

int kind#

Caller-defined token kind (e.g. an enum value).

std::string_view lexeme#

The matched text, viewing into the source.

position start#

Position of the token’s first byte.

std::size_t mode_id = {0}#

The mode this token was lexed in (0 = the default/root mode).

struct position#

A location in the source text.

Columns are counted within the line in the unit the lexer’s column_unit names, bytes by default (REAL’s byte-level UTF-8 model: a multibyte code point spans several columns), or code points, or UTF-16 code units. Lines and columns are 1-based; a line ends at \n only (a lone \r does not end one). offset is a 0-based byte index from the start of the source, whatever the column unit.

Public Members

std::size_t offset#

0-based byte offset from the start of the source.

std::size_t line#

1-based line number.

std::size_t column#

1-based column within the line, in the lexer’s column_unit.

class lex_error : public std::runtime_error#

Thrown when no rule matches at a position (a lexical error).

Carries the position of the offending byte so a caller can report it.

Public Functions

inline lex_error(const std::string &message, position where)#

Builds the error.

Parameters:
  • message – [in] Human-readable cause.

  • where – [in] Position of the byte that no rule could match.

inline lex_error(const std::string &message, position where, std::vector<token> decided)#

Builds the error a stream raises after deciding tokens in the same call.

Parameters:
  • message – [in] Human-readable cause.

  • where – [in] Position of the byte that no rule could match.

  • decided – [in] The tokens the call decided before it, in order.

inline position where() const noexcept#

Returns the position of the unmatched byte.

inline const std::vector<token> &decided() const noexcept#

The tokens a token_stream call decided before the error, in order: with the tokens its earlier calls returned, every token before the error, wherever the text was cut. Empty from lexer::tokenize.

Returns:

The tokens; a lexeme views the stream’s buffer, valid until its next call.

Layout (scilex/layout.hpp)#

inline std::vector<token> scilex::layout(std::span<const token> tokens, const std::vector<bool> &mode_significant = {})#

Rewrites tokens with NEWLINE / INDENT / DEDENT inserted.

Parameters:
  • tokens – [in] An end-of-input-terminated token sequence.

  • mode_significant – [in] Per-mode-id significance policy (Layout Awareness Level A): index by a token’s mode_id; true (or a mode-id beyond the vector) means the token shapes layout, false means it is passed through without affecting indentation. An empty vector (the default) means every token is significant — byte-for-byte the positional pass. (A std::vector<bool> rather than a std::span<const bool>: the bit-packed vector<bool> cannot be viewed as a contiguous span of bool.)

Throws:

layout_error – If a line dedents to an indentation that no open block used.

Returns:

The layout-aware token sequence (still end-of-input-terminated).

inline std::vector<token> scilex::layout(std::span<const token> tokens, std::string_view source, tab_policy tabs, const std::vector<bool> &mode_significant = {})#

Rewrites tokens with NEWLINE / INDENT / DEDENT inserted, measuring indentation in source under tabs.

Parameters:
  • tokens – [in] An end-of-input-terminated token sequence lexed from source.

  • source – [in] The text tokens were lexed from; a token’s start.offset indexes it.

  • tabs – [in] How indentation is measured (tab_policy).

  • mode_significant – [in] As in the overload without a source.

Throws:
  • layout_error – If a line dedents to an indentation that no open block used, or (under tab_policy::python) mixes tabs and spaces so that its level is ambiguous.

  • std::invalid_argument – If a token’s offset lies beyond source.

Returns:

The layout-aware token sequence (still end-of-input-terminated).

enum class scilex::tab_policy : std::uint8_t#

How the overload of layout() that takes the source measures indentation.

Values:

enumerator columns#

A tab is one column, like a space; mixed tabs and spaces are not policed (the default pass).

enumerator python#

CPython’s rule. A line’s indentation is measured twice over the blanks before its first token: with tabs advancing to the next multiple of 8, and with every tab counted as 1. Levels compare by the first measure; when the second does not order the same way — a deeper, equal or shallower line by one measure and not by the other — the line is refused with “inconsistent use of tabs and spaces in indentation”, the message CPython’s TabError carries. A form feed resets both measures, as in CPython.

class layout_error : public std::runtime_error#

Thrown when a line’s indentation matches no enclosing level.

Public Functions

inline layout_error(const std::string &message, position where)#

Builds the error.

Parameters:
  • message – [in] Cause.

  • where – [in] Position.

inline position where() const noexcept#

Returns the position of the offending line.

Grammars (scilex/grammar.hpp)#

inline grammar scilex::parse_grammar(std::string_view text, const std::string &origin = "<string>")#

Parses a .lex grammar from text (see the file documentation for the format).

Never returns a half-built grammar: the first malformed line throws.

Parameters:
  • text – [in] The grammar.

  • origin – [in] What errors name as the grammar’s origin (a path, or a label).

Throws:

grammar_error – On a malformed line, an invalid pattern (at its column in the line), or a grammar with no rules.

Returns:

The rules and their names.

inline grammar scilex::load_grammar(const std::string &path)#

Reads and parses the .lex grammar at path.

Parameters:

path – [in] The grammar file.

Throws:

grammar_error – If the file cannot be read (line 0) or is malformed (see parse_grammar).

Returns:

The rules and their names.

struct grammar#

A parsed grammar: the rules in order, and each rule’s name (index = the rule’s kind).

Public Functions

inline std::string_view name(int kind) const noexcept#

The name of kind, or "?" for a kind no rule of this grammar has (a reserved kind such as error).

Parameters:

kind – [in] A token kind.

Returns:

Its name.

Public Members

std::vector<rule> rules#

The rules, kind i at index i.

std::vector<std::string> names#

The name of kind i.

class grammar_error : public std::runtime_error#

Thrown for a malformed .lex grammar: where (origin, 1-based line, and 1-based byte column in the line when the cause has one) and why.

what() reads origin:line[:column]: cause, the form compilers and editors recognize, or origin: cause for an error with no line (an unreadable file, a grammar with no rules).

Public Functions

inline grammar_error(const std::string &origin, std::size_t line, std::size_t column, const std::string &cause)#

Builds the error.

Parameters:
  • origin – [in] Where the grammar came from (a path, or a label such as <string>).

  • line – [in] The 1-based line of the offending rule.

  • column – [in] The 1-based byte column in that line, or 0 when the cause has none.

  • cause – [in] What is wrong, without the location.

inline std::size_t line() const noexcept#

The 1-based line of the offending rule (0 for a grammar-wide error such as no rules).

Returns:

The line.

inline std::size_t column() const noexcept#

The 1-based byte column in the line, or 0 when the cause has none.

Returns:

The column.

inline const std::string &cause() const noexcept#

What is wrong, without the location.

Returns:

The cause.