Python#

class scilex.Lexer(rules, insignificant_modes=(), dfa_modes=(), errors='raise', columns='bytes', dfa='auto')#

A compiled, reusable set of token rules.

Parameters:
  • rules (iterable) – Items (kind, pattern[, skip[, in_mode[, action]]]); kind is an int, pattern a REAL regex string, skip a bool (default False) — when true, matches are consumed but not emitted. in_mode is a sequence of mode names the rule is active in (empty, the default, means the default mode only); action drives the per-scan mode stack and is None or one of ("push", mode) / ("set", mode) / ("pop",). These last two are SciLex’s modes (contextual lexing); a plain (kind, pattern, skip) rule needs neither.

  • insignificant_modes (iterable) – Mode names whose tokens carry no layout structure (Layout Awareness Level A — see layout()): code spanning lines in such a mode is treated as continuation. Each must be a mode the rules use.

  • dfa_modes (iterable) – Mode names to accelerate with a DFA fast path (one DFA pass replaces the per-rule dispatch) when dfa="requested"; under the default dfa="auto" every mode is tried and this adds nothing. Best-effort and invisible: a mode whose rules need an assertion no DFA can represent, or whose DFA would change an answer (a rule whose match is not its longest match, such as as|assert), silently stays on the regular engine — see dfa_modes_active. The decision is exact, so the token stream is identical either way.

  • errors (str) –

    What to do at a byte no rule can lex. The default "raise" is unchanged — it raises error at the first unlexable byte, exactly as before. "token" opts into recovery: a maximal run of unlexable bytes is emitted as one token of reserved kind ERROR (its lexeme the exact bytes) and lexing resumes, so a whole document is tokenized in one pass.

    Recovery cost is the cost of a no-match in your grammar: at each byte of an error run a first-byte pre-filter skips positions no rule can begin (usually O(1) per byte), attempting a match only where a rule might start. A rule that scans far before failing — an unanchored, greedy pattern with no distinguishing leading byte — pays that scan at every position of a long error run (on the order of 200 µs per position over an 8 KB run in the worst case). Prefer rules with a definite leading byte if recovery speed on hostile input matters.

  • columns (str) – The unit each token’s column is counted in. The default "bytes" is unchanged (column == byte offset within the line + 1). "codepoints" counts Unicode scalar values; "utf16" counts UTF-16 code units (an astral codepoint is 2 — the unit an LSP client expects). A malformed byte counts as one unit in every mode. The unit is not stored on a Position — read it back from column_unit.

  • dfa (str) – Which modes are tried for DFA acceleration. "auto" (the default): every mode, each keeping its DFA only where it reproduces the per-rule munch exactly. "requested": only the modes in dfa_modes (none, if empty).

Raises:
  • error – If a pattern is an invalid regex, a transition targets an empty mode, or dfa_modes names an unknown mode.

  • ValueError – If insignificant_modes names a mode the rules do not use.

property rules#

(kind, pattern, skip), or (kind, pattern, skip, in_mode, action) for rules that use modes.

Type:

The normalized rules

property insignificant_modes#

The mode names whose tokens carry no layout structure (Level A).

property dfa_modes#

The mode names requested for DFA acceleration (see dfa_modes_active).

property dfa_modes_active#

The mode names with at least one rule on a DFA.

In an accelerated mode, a rule the DFA cannot take – it needs an assertion no DFA can represent, or its match is not its longest match – stays on the regular engine beside the DFA (see pike_rules()); a mode where no rule can go on one is absent. The tokens are the same either way.

pike_rules(mode='default')#

The indices (into rules) of the rules of mode that run on the regular engine.

Every rule of a mode with no DFA; in an accelerated mode, only the rules its DFA cannot take.

Raises:

error – If mode is not a mode the rules use.

end_of(source, token)#

Where token ends in source: the Position just past its last byte, counted in this lexer’s column_unit – where the scan stood after the token.

A token carries its start, not its end, so the stream costs nothing more for the callers that never ask. A zero-width token (NEWLINE, INDENT, DEDENT from layout, END_OF_INPUT) ends where it starts.

Parameters:
  • source (str | bytes) – The text token was lexed from.

  • token (Token) – A token this lexer produced from source.

Returns:

The position after token.

Return type:

Position

Raises:

ValueError – If token’s lexeme is not the text of source at its offset.

property column_unit#

"bytes" (the default), "codepoints", or "utf16".

A Position does not carry its unit — this is where the lexer declares it, a deliberate one-field trade-off. Read it when a consumer (an editor, an LSP server) must interpret a column.

Type:

The unit each token’s position.column is counted in

layout(tokens, source=None, tabs='columns')#

Insert NEWLINE / INDENT / DEDENT from indentation, mode-aware.

Uses this lexer’s insignificant_modes (Layout Awareness Level A): a token in an insignificant mode is passed through without shaping indentation, so multi-line brackets / flow collections read as continuation.

Parameters:
  • tokens (iterable[Token]) – An end-of-input-terminated token stream (tokenize with eof=True).

  • source (str | None) – The text tokens were lexed from; needed by tabs="python".

  • tabs (str) – "columns" or "python" — see Layout.apply().

Returns:

The layout-aware tokens (still END_OF_INPUT-terminated).

Return type:

list[Token]

Raises:

error – On a line that dedents to an unknown indentation, or mixes tabs and spaces ambiguously under tabs="python" (.position).

tokenize(text, eof=False)#

Tokenize text eagerly into a list.

Parameters:
  • text (str | bytes) – The source to tokenize; each Token’s lexeme is a str when text is str, bytes when it is bytes.

  • eof (bool) – Append a terminal END_OF_INPUT token at the end.

Returns:

Emitted tokens in source order (skip matches omitted).

Return type:

list[Token]

Raises:

error – If some position is matched by no rule (with .position).

scan(text, eof=False)#

Lazily scan text, yielding one Token at a time.

Unlike tokenize(), nothing but the current token is held — the parser-friendly access pattern. The returned generator runs the lexer on demand; a lexical error surfaces while iterating, only after every token before it has been yielded.

scan holds the GIL for each one-token step (the parser-friendly access pattern); for multi-threaded throughput use tokenize(), which releases the GIL around the scan of inputs of 4 KB or more.

Parameters:
  • text (str | bytes) – The source to scan; each Token’s lexeme is a str when text is str, bytes when it is bytes.

  • eof (bool) – Yield a terminal END_OF_INPUT token at the end.

Yields:

Token – The next token in source order (skip matches omitted).

Raises:

error – If some position is matched by no rule (with .position).

stream()#

A TokenStream over this lexer: text fed in pieces, lexed as it arrives.

Returns:

A stream at the start of the text.

Return type:

TokenStream

class scilex.Token#
column#

position.column — 1-based column, in the lexer’s column unit

kind#

the token kind (int)

lexeme#

the matched text (str or bytes)

line#

position.line — 1-based line

mode#

the name of the mode the token was lexed in

offset#

position.offset — 0-based byte offset

position#

the Position of the first byte

class scilex.Position#
column#

1-based column of the position, in the lexer’s column unit (bytes by default)

line#

1-based line of the position

offset#

0-based byte offset of the position

class scilex.Layout(insignificant_modes=())#

Inserts NEWLINE / INDENT / DEDENT tokens from a token stream’s indentation.

Indentation-significant languages (Python-like, e.g. SciLang) read structure from leading whitespace. Working purely from token positions, apply() rewrites a flat token stream into a layout-aware one: a NEWLINE at each logical line end, and INDENT / DEDENT where the leading (byte) column of a line’s first token changes. Blank and comment-only lines carry no token, so they add no structure.

With insignificant_modes (Layout Awareness Level A), a token whose Token.mode is listed is passed through without shaping indentation, so a multi-line bracket or flow collection reads as line continuation. The default (none) is the positional pass. Lexer.layout() wires a lexer’s own modes.

The reserved kinds are exposed as newline_kind, indent_kind and dedent_kind.

apply(tokens, source=None, tabs='columns')#

Rewrite tokens with NEWLINE/INDENT/DEDENT inserted.

Parameters:
  • tokens (iterable[Token]) – An end-of-input-terminated token stream — tokenize with eof=True (the terminal END_OF_INPUT is preserved). Each token’s kind, position and mode are read.

  • source (str | None) – The text tokens were lexed from. Only tabs="python" reads it.

  • tabs (str) – How indentation is measured. "columns" (the default): a tab is one column, like a space. "python": CPython’s rule — tabs advance to the next multiple of 8, and a line whose level would differ with tabs counted as 1 is refused with “inconsistent use of tabs and spaces in indentation”; a form feed resets the measure. Needs source.

Returns:

The layout-aware tokens (still END_OF_INPUT-terminated).

Return type:

list[Token]

Raises:
  • error – On a line that dedents to an indentation no open block used, or (tabs="python") mixes tabs and spaces so that its level is ambiguous, carrying .position (no .context snippet).

  • ValueError – On an unknown tabs, tabs="python" without source, or a token whose offset lies beyond source.

scilex.tokenize(rules, text, eof=False)#

Compile rules and tokenize text eagerly in one call.

Parameters:
  • rules (iterable) – See Lexer.

  • text (str) – The source to tokenize.

  • eof (bool) – Append a terminal END_OF_INPUT token.

Returns:

Emitted tokens in source order.

Return type:

list[Token]

scilex.scan(rules, text, eof=False)#

Compile rules and lazily scan text in one call.

Parameters:
  • rules (iterable) – See Lexer.

  • text (str) – The source to scan.

  • eof (bool) – Yield a terminal END_OF_INPUT token.

Yields:

Token – The next token in source order.

scilex.layout(tokens, insignificant_modes=(), source=None, tabs='columns')#

Insert NEWLINE/INDENT/DEDENT into tokens (see Layout.apply()).

Parameters:
  • tokens (iterable[Token]) – An end-of-input-terminated token stream.

  • insignificant_modes (iterable) – Mode names that carry no layout structure (Layout Awareness Level A).

  • source (str | None) – The text tokens were lexed from (for tabs="python").

  • tabs (str) – "columns" or "python" (see Layout.apply()).

Returns:

The layout-aware tokens.

Return type:

list[Token]

Grammars#

scilex.parse_grammar(text, origin='<string>')#

Parse a .lex grammar from text (see Grammar for the format).

Parameters:
  • text (str) – The grammar.

  • origin (str) – What errors name as its origin (a path, or a label).

Returns:

The rules and their names.

Return type:

Grammar

Raises:

GrammarError – On a malformed line, an invalid pattern (at its column), or no rules.

scilex.load_grammar(path)#

Read and parse the .lex grammar at path (see parse_grammar()).

Raises:

GrammarError – If the file is malformed; OSError if it cannot be read.

class scilex.Grammar(rules, names)#

A parsed .lex grammar: rules in the form Lexer takes, and names, the name of each rule, indexed by its kind (a rule’s kind is its position in the file).

Build one with parse_grammar() or load_grammar(); the format is the CLI’s (one rule per line, name<TAB>pattern[<TAB>options], options skip, in=m1,m2 and one of push=m, set=m, pop).

name(kind)#

The name of kind, or "?" for a kind no rule of this grammar has (a reserved kind).

lexer(insignificant_modes=(), dfa_modes=(), errors='raise', columns='bytes', dfa='auto')#

A Lexer over these rules, with Lexer’s options.

Errors#

exception scilex.error#
exception scilex.LexError#

Input no rule can lex: a byte no rule matches, a zero-length winning match, a pop at the root mode or a push past the mode stack’s bound, or input ending inside a pushed mode. Carries .position (and .context where the source is known). Raised by a stream’s feed or finish, it also carries .tokens: the tokens that call decided before the error, so the tokens of every call before it and these are all the tokens before the error, wherever the text was cut. A subclass of error.

exception scilex.LayoutError#

Indentation Layout.apply() cannot lay out: a dedent to no open level, or (with tabs="python") tabs and spaces mixed so that a level is ambiguous. Carries .position. A subclass of error.

exception scilex.GrammarError#

A malformed .lex grammar (parse_grammar(), load_grammar()). Carries .line (1-based; 0 for a grammar-wide error such as no rules), .column (1-based byte column in the line; 0 when the cause has none) and .cause (the message without its location). A subclass of error.

Constants and helpers#

END_OF_INPUT, NEWLINE, INDENT, DEDENT and ERROR are the reserved token kinds.

scilex.get_include()#

Return the directory to add to a C++ include path for SciLex’s headers.

#include <scilex/scilex.hpp> resolves against this directory. SciLex is header-only and shipped inside the installed package, so a project can compile against it located through its Python install. Note that SciLex’s headers include REAL’s, so add real.get_include() as well.

Returns:

Absolute path to SciLex’s include directory.

Return type:

str

scilex.get_config()#

Return metadata for embedding the C++ library.

Returns:

version (str), include (str, see get_include()), and cxx_standard (str).

Return type:

dict

scilex.real_version() → str#

The REAL version this compiled extension was built against (real/version.hpp), so a stale build or install can be detected by comparing it to the pinned real-regex version.