Python#
- class scilex.Lexer(rules, insignificant_modes=(), dfa_modes=(), errors='raise', columns='bytes', dfa='auto')#
A compiled, reusable set of token rules.
- Parameters:
rules (iterable) – Items
(kind, pattern[, skip[, in_mode[, action]]]);kindis an int,patterna REAL regex string,skipa bool (defaultFalse) — when true, matches are consumed but not emitted.in_modeis a sequence of mode names the rule is active in (empty, the default, means the default mode only);actiondrives the per-scan mode stack and isNoneor one of("push", mode)/("set", mode)/("pop",). These last two are SciLex’s modes (contextual lexing); a plain(kind, pattern, skip)rule needs neither.insignificant_modes (iterable) – Mode names whose tokens carry no layout structure (Layout Awareness Level A — see
layout()): code spanning lines in such a mode is treated as continuation. Each must be a mode the rules use.dfa_modes (iterable) – Mode names to accelerate with a DFA fast path (one DFA pass replaces the per-rule dispatch) when
dfa="requested"; under the defaultdfa="auto"every mode is tried and this adds nothing. Best-effort and invisible: a mode whose rules need an assertion no DFA can represent, or whose DFA would change an answer (a rule whose match is not its longest match, such asas|assert), silently stays on the regular engine — seedfa_modes_active. The decision is exact, so the token stream is identical either way.errors (str) –
What to do at a byte no rule can lex. The default
"raise"is unchanged — it raiseserrorat the first unlexable byte, exactly as before."token"opts into recovery: a maximal run of unlexable bytes is emitted as one token of reserved kindERROR(itslexemethe exact bytes) and lexing resumes, so a whole document is tokenized in one pass.Recovery cost is the cost of a no-match in your grammar: at each byte of an error run a first-byte pre-filter skips positions no rule can begin (usually O(1) per byte), attempting a match only where a rule might start. A rule that scans far before failing — an unanchored, greedy pattern with no distinguishing leading byte — pays that scan at every position of a long error run (on the order of 200 µs per position over an 8 KB run in the worst case). Prefer rules with a definite leading byte if recovery speed on hostile input matters.
columns (str) – The unit each token’s column is counted in. The default
"bytes"is unchanged (column == byte offset within the line + 1)."codepoints"counts Unicode scalar values;"utf16"counts UTF-16 code units (an astral codepoint is 2 — the unit an LSP client expects). A malformed byte counts as one unit in every mode. The unit is not stored on aPosition— read it back fromcolumn_unit.dfa (str) – Which modes are tried for DFA acceleration.
"auto"(the default): every mode, each keeping its DFA only where it reproduces the per-rule munch exactly."requested": only the modes indfa_modes(none, if empty).
- Raises:
error – If a pattern is an invalid regex, a transition targets an empty mode, or
dfa_modesnames an unknown mode.ValueError – If
insignificant_modesnames a mode the rules do not use.
- property rules#
(kind, pattern, skip), or(kind, pattern, skip, in_mode, action)for rules that use modes.- Type:
The normalized rules
- property insignificant_modes#
The mode names whose tokens carry no layout structure (Level A).
- property dfa_modes#
The mode names requested for DFA acceleration (see
dfa_modes_active).
- property dfa_modes_active#
The mode names with at least one rule on a DFA.
In an accelerated mode, a rule the DFA cannot take – it needs an assertion no DFA can represent, or its match is not its longest match – stays on the regular engine beside the DFA (see
pike_rules()); a mode where no rule can go on one is absent. The tokens are the same either way.
- pike_rules(mode='default')#
The indices (into
rules) of the rules ofmodethat run on the regular engine.Every rule of a mode with no DFA; in an accelerated mode, only the rules its DFA cannot take.
- Raises:
error – If
modeis not a mode the rules use.
- end_of(source, token)#
Where
tokenends insource: thePositionjust past its last byte, counted in this lexer’scolumn_unit– where the scan stood after the token.A token carries its start, not its end, so the stream costs nothing more for the callers that never ask. A zero-width token (NEWLINE, INDENT, DEDENT from layout, END_OF_INPUT) ends where it starts.
- property column_unit#
"bytes"(the default),"codepoints", or"utf16".A
Positiondoes not carry its unit — this is where the lexer declares it, a deliberate one-field trade-off. Read it when a consumer (an editor, an LSP server) must interpret a column.- Type:
The unit each token’s
position.columnis counted in
- layout(tokens, source=None, tabs='columns')#
Insert NEWLINE / INDENT / DEDENT from indentation, mode-aware.
Uses this lexer’s
insignificant_modes(Layout Awareness Level A): a token in an insignificant mode is passed through without shaping indentation, so multi-line brackets / flow collections read as continuation.- Parameters:
tokens (iterable[Token]) – An end-of-input-terminated token stream (tokenize with
eof=True).source (str | None) – The text
tokenswere lexed from; needed bytabs="python".tabs (str) –
"columns"or"python"— seeLayout.apply().
- Returns:
The layout-aware tokens (still END_OF_INPUT-terminated).
- Return type:
list[Token]
- Raises:
error – On a line that dedents to an unknown indentation, or mixes tabs and spaces ambiguously under
tabs="python"(.position).
- tokenize(text, eof=False)#
Tokenize
texteagerly into a list.- Parameters:
text (str | bytes) – The source to tokenize; each
Token’s lexeme is astrwhentextisstr,byteswhen it isbytes.eof (bool) – Append a terminal
END_OF_INPUTtoken at the end.
- Returns:
Emitted tokens in source order (skip matches omitted).
- Return type:
list[Token]
- Raises:
error – If some position is matched by no rule (with
.position).
- scan(text, eof=False)#
Lazily scan
text, yielding oneTokenat a time.Unlike
tokenize(), nothing but the current token is held — the parser-friendly access pattern. The returned generator runs the lexer on demand; a lexical error surfaces while iterating, only after every token before it has been yielded.scanholds the GIL for each one-token step (the parser-friendly access pattern); for multi-threaded throughput usetokenize(), which releases the GIL around the scan of inputs of 4 KB or more.- Parameters:
text (str | bytes) – The source to scan; each
Token’s lexeme is astrwhentextisstr,byteswhen it isbytes.eof (bool) – Yield a terminal
END_OF_INPUTtoken at the end.
- Yields:
Token – The next token in source order (skip matches omitted).
- Raises:
error – If some position is matched by no rule (with
.position).
- stream()#
A
TokenStreamover this lexer: text fed in pieces, lexed as it arrives.- Returns:
A stream at the start of the text.
- Return type:
TokenStream
- class scilex.Token#
- column#
position.column — 1-based column, in the lexer’s column unit
- kind#
the token kind (int)
- lexeme#
the matched text (str or bytes)
- line#
position.line — 1-based line
- mode#
the name of the mode the token was lexed in
- offset#
position.offset — 0-based byte offset
- position#
the Position of the first byte
- class scilex.Position#
- column#
1-based column of the position, in the lexer’s column unit (bytes by default)
- line#
1-based line of the position
- offset#
0-based byte offset of the position
- class scilex.Layout(insignificant_modes=())#
Inserts NEWLINE / INDENT / DEDENT tokens from a token stream’s indentation.
Indentation-significant languages (Python-like, e.g. SciLang) read structure from leading whitespace. Working purely from token positions,
apply()rewrites a flat token stream into a layout-aware one: aNEWLINEat each logical line end, andINDENT/DEDENTwhere the leading (byte) column of a line’s first token changes. Blank and comment-only lines carry no token, so they add no structure.With
insignificant_modes(Layout Awareness Level A), a token whoseToken.modeis listed is passed through without shaping indentation, so a multi-line bracket or flow collection reads as line continuation. The default (none) is the positional pass.Lexer.layout()wires a lexer’s own modes.The reserved kinds are exposed as
newline_kind,indent_kindanddedent_kind.- apply(tokens, source=None, tabs='columns')#
Rewrite
tokenswith NEWLINE/INDENT/DEDENT inserted.- Parameters:
tokens (iterable[Token]) – An end-of-input-terminated token stream — tokenize with
eof=True(the terminalEND_OF_INPUTis preserved). Each token’skind, position andmodeare read.source (str | None) – The text
tokenswere lexed from. Onlytabs="python"reads it.tabs (str) – How indentation is measured.
"columns"(the default): a tab is one column, like a space."python": CPython’s rule — tabs advance to the next multiple of 8, and a line whose level would differ with tabs counted as 1 is refused with “inconsistent use of tabs and spaces in indentation”; a form feed resets the measure. Needssource.
- Returns:
The layout-aware tokens (still END_OF_INPUT-terminated).
- Return type:
list[Token]
- Raises:
error – On a line that dedents to an indentation no open block used, or (
tabs="python") mixes tabs and spaces so that its level is ambiguous, carrying.position(no.contextsnippet).ValueError – On an unknown
tabs,tabs="python"withoutsource, or a token whose offset lies beyondsource.
- scilex.tokenize(rules, text, eof=False)#
Compile
rulesand tokenizetexteagerly in one call.
- scilex.scan(rules, text, eof=False)#
Compile
rulesand lazily scantextin one call.- Parameters:
rules (iterable) – See
Lexer.text (str) – The source to scan.
eof (bool) – Yield a terminal
END_OF_INPUTtoken.
- Yields:
Token – The next token in source order.
- scilex.layout(tokens, insignificant_modes=(), source=None, tabs='columns')#
Insert NEWLINE/INDENT/DEDENT into
tokens(seeLayout.apply()).- Parameters:
tokens (iterable[Token]) – An end-of-input-terminated token stream.
insignificant_modes (iterable) – Mode names that carry no layout structure (Layout Awareness Level A).
source (str | None) – The text
tokenswere lexed from (fortabs="python").tabs (str) –
"columns"or"python"(seeLayout.apply()).
- Returns:
The layout-aware tokens.
- Return type:
list[Token]
Grammars#
- scilex.parse_grammar(text, origin='<string>')#
Parse a
.lexgrammar fromtext(seeGrammarfor the format).- Parameters:
text (str) – The grammar.
origin (str) – What errors name as its origin (a path, or a label).
- Returns:
The rules and their names.
- Return type:
- Raises:
GrammarError – On a malformed line, an invalid pattern (at its column), or no rules.
- scilex.load_grammar(path)#
Read and parse the
.lexgrammar atpath(seeparse_grammar()).- Raises:
GrammarError – If the file is malformed;
OSErrorif it cannot be read.
- class scilex.Grammar(rules, names)#
A parsed
.lexgrammar:rulesin the formLexertakes, andnames, the name of each rule, indexed by its kind (a rule’s kind is its position in the file).Build one with
parse_grammar()orload_grammar(); the format is the CLI’s (one rule per line,name<TAB>pattern[<TAB>options], optionsskip,in=m1,m2and one ofpush=m,set=m,pop).- name(kind)#
The name of
kind, or"?"for a kind no rule of this grammar has (a reserved kind).
Errors#
- exception scilex.error#
- exception scilex.LexError#
Input no rule can lex: a byte no rule matches, a zero-length winning match, a pop at the root mode or a push past the mode stack’s bound, or input ending inside a pushed mode. Carries
.position(and.contextwhere the source is known). Raised by a stream’sfeedorfinish, it also carries.tokens: the tokens that call decided before the error, so the tokens of every call before it and these are all the tokens before the error, wherever the text was cut. A subclass oferror.
- exception scilex.LayoutError#
Indentation
Layout.apply()cannot lay out: a dedent to no open level, or (withtabs="python") tabs and spaces mixed so that a level is ambiguous. Carries.position. A subclass oferror.
- exception scilex.GrammarError#
A malformed
.lexgrammar (parse_grammar(),load_grammar()). Carries.line(1-based; 0 for a grammar-wide error such as no rules),.column(1-based byte column in the line; 0 when the cause has none) and.cause(the message without its location). A subclass oferror.
Constants and helpers#
END_OF_INPUT, NEWLINE, INDENT, DEDENT and ERROR are the reserved token kinds.
- scilex.get_include()#
Return the directory to add to a C++ include path for SciLex’s headers.
#include <scilex/scilex.hpp>resolves against this directory. SciLex is header-only and shipped inside the installed package, so a project can compile against it located through its Python install. Note that SciLex’s headers include REAL’s, so addreal.get_include()as well.- Returns:
Absolute path to SciLex’s include directory.
- Return type:
str
- scilex.get_config()#
Return metadata for embedding the C++ library.
- Returns:
version(str),include(str, seeget_include()), andcxx_standard(str).- Return type:
dict
- scilex.real_version() str#
The REAL version this compiled extension was built against (real/version.hpp), so a stale build or install can be detected by comparing it to the pinned real-regex version.