Skip to content

how transpilation works

basedpython transpiles .by source files into standard Python. the pipeline is deterministic — a single AST-rewrite pass followed by a few text normalization phases and a final verification step, no convergence loop

pipeline

source (.by)
  │
  ├─ phase 0  AST rewrite (ast_driver::run_against_source)
  │     ├─ strip use-site variance keywords (`out`/`in`) up front
  │     ├─ parse via ruff_python_parser (the unified parser accepts `.by` syntax)
  │     ├─ build a SemanticModel — the project db when available (cross-module
  │     │  type info), else a single-file in-memory db. a pre-pass that rewrote
  │     │  the source doesn't cost the project: the db is rebuilt over the
  │     │  rewritten text (see "the phase-0 database" below)
  │     ├─ run AstPasses: mutate the AST in place (coalesce, cast, typeof,
  │     │  sentinel, mutable-defaults, …)
  │     ├─ run TypeAwarePasses: read the SemanticModel and emit text edits
  │     │  (intersection, callable, generics, literal-types, anon-NT, …)
  │     └─ splice it together: re-emit the nodes the AST passes changed, apply
  │        text edits (ruff-style first-wins overlap skip, composed inside the
  │        nodes both touch), emit hoisted class defs, prepend
  │        required imports and the runtime helpers the emitted code calls
  │        (see "the runtime" below), append `__all__` epilogue
  │
  ├─ phase 1  lowering preamble
  │     └─ optionally prepend `from __future__ import annotations`
  │        (`inject_future_annotations`, off by default)
  │
  ├─ phase 2  import redirect
  │     └─ rewrite `from typing import X` to `from typing_extensions import X`
  │        when X is not yet stdlib at the configured min Python version
  │
  ├─ phase 2b  anon-named-tuple cleanup
  │     └─ re-run `anon_named_tuple` on the post-lowering output to catch anon-NT
  │        spans copied verbatim by other transforms (e.g. the PEP-695 polyfill).
  │        bounded to a few iterations
  │
  ├─ phase 2c  lazy-import marking
  │     └─ lower imports to the `lazy` keyword (3.15+) or a runtime polyfill
  │
  ├─ phase 2d  version polyfill
  │     └─ rewrite syntax the target python cannot parse. today that is the
  │        `match` statement, lowered to an `if`/`elif` chain for a target below
  │        3.10. it runs over the *finished* python rather than over `.by`, so a
  │        `match` an earlier lowering generated is lowered by the same code as
  │        one the author wrote
  │
  └─ phase 3  syntax verification
        ├─ parse the final output as `.py` — any parse error aborts with a
        │  source-annotated diagnostic (the span is mapped back to `.by`)
        ├─ scan the AST for leftover basedpython-only flags
        │  (`is_anon_named_tuple`, `is_anon_named_tuple_value`, `is_typeof`).
        │  a leftover flag means a transform failed to lower its construct; the
        │  pipeline aborts rather than emit syntactically-valid-but-wrong Python
        ├─ reject a call to one of the transpiler's own runtime helpers that the
        │  module never got. a transform emits the call and records the need in
        │  two different places, and forgetting the second half produces python
        │  that parses, checks, and raises `NameError` the first time the lowered
        │  line runs
        └─ parse it again *as the target version* and report any construct that
           version cannot parse. the first check asks whether the output is
           python at all; this one asks whether it is python the declared floor
           can run, so syntax no polyfill covers is a diagnostic instead of a
           `SyntaxError` at import time in generated code. which kinds of
           syntax a lowering removes is stated once, by
           `UnsupportedSyntaxErrorKind::is_lowered_by_basedpython`; the parser
           holds a `.by` file to its target for every other kind, so a checker
           reports what this check would refuse. a basedpython form written as
           syntax of its own is a kind in that list too: a destructuring
           pattern is written as a `match`, and is held to where one runs

entry points in crates/by_transforms/src/lib.rs:

  • transpile(source, config) -> Result<String, String> — single-file (stdin, tests); type-aware passes see only this file
  • transpile_typed(db, file, config, rebuild) — uses an existing project db so type-aware passes resolve cross-module types (the CLI path for by transpile, by build and by run)
  • transpile_typed_with_map(db, file, config, rebuild) — also returns a line table for traceback rewriting and diagnostic mapping

the runtime

the emitted python calls helpers of its own: _lazy_module for a lowered import, _parametric_is for a runtime type test, _by_loop_bind for a closure that captures a loop binding by value. they live in crates/by_transforms/src/runtime/_by_runtime.py, and runtime.rs slices them out by name

a transform names the helpers the code it emits calls, through the typed constants in runtime.rs. what each of those calls in turn is read out of its body, so a helper brings the rest of what it needs along

Config::runtime_module decides how a module gets them:

  • by build, by run and by compile stage a tree of their own. they write _by_runtime.py into each package they stage — a directory with an __init__ — and each module imports what it calls from its package's copy
  • a module at a module root gets the definitions pasted in. a copy at the root would be a top-level module, which a second basedpython wheel built by another version overwrites on install
  • so does a module in a directory that is not a package, a scripts/ folder say, which has no import that works both when it is run and when it is imported
  • by transpile <file> and the language server's by/transpile answer with one module's text, and by transpile <dir> writes into the source tree itself, where a file with no .by beside it would read as a module the author wrote. all three paste the definitions in

both renderings are slices of the same text. the lazy-import pass leaves the runtime import eager, since a helper reached through its proxy would be a proxy call on every use

_by_runtime.py is excluded from this repository's own ruff configuration. reformatting it rewrites transpiled output, and the isort rule would put a from __future__ import annotations at the top of every module that gets a helper pasted into it

names the lowerings write

a lowering writes names of its own into the module: a temporary, a TypeVar the generics polyfill declares, and the typing constructs it spells a type with — Union, Callable, Literal, final, functools.partial and the rest. each is chosen through WrittenNames (transforms/repeated_underscore.rs), so the names cannot collide with the module's own:

  • a binding the lowering writes (fresh) takes a name the module spells nowhere, and none a runtime helper is bound under
  • a typing name (imported, import_from) keeps its own spelling unless the module binds that name to something else, and is imported under a fresh name when it does. a module that binds it only by importing the very same thing — from typing import Literal beside a lowered 1 | 2 — reads the lowering's already. a star import binds names unseen, so a module with one gets a fresh name for every typing name
  • a builtin (builtin) is the same: isinstance, staticmethod, type keep their names unless the module binds them anywhere — a parameter named type counts — and are then imported from builtins under a fresh name, which the driver binds once it sees the output read it
  • a module-level private symbol is emitted under _name unless the module, a runtime helper or another lowering already has that name, and then under a fresh one (PrivateNames). every other fresh name stays clear of the ones it chose
Union = "mine"
x = list[int?]
# generated python, for 3.9
from typing import Union as Union2
Union = "mine"
x = list[Union2[int, None]]

a keyword the parser stands up as a name — the final of final class — binds nothing, so a module that writes one still gets @final

the phase-0 database

three pre-passes run before phase 0 and can rewrite the source: erased-union reification, context-sensitive name qualification, and enum lowering. phase 0 re-parses its input and binds inferred_type to those exact node identities, so its database has to hold the rewritten text, not the project file's

that would leave phase 0 with a single-file in-memory db — no search paths, no sibling modules — and every TypeAwarePass silently declining to lower, so an enum declared anywhere in a file would break a lookup elsewhere in it. instead the caller supplies a RebuildProject: a way to build a second database over the same project, into which the rewritten source is served for this one file via File::source_text_override. the project's metadata, search paths and sibling files stay; only the file being transpiled reads differently

the rebuilt database must be a new one rather than a clone — salsa handles cloned from one database share storage, so the override would be visible through the caller's own db. a caller with no project (a bare unit test) passes None and gets the single-file db: correct, just blind past the file

passes

ast_driver runs two kinds of pass against the parsed module:

  • AstPass — mutates the AST in place via the Transformer protocol. the driver tracks which top-level statements changed, and re-emits only the nodes in them the pass changed, through ruff_python_codegen (basedpython mode). a node it made carries no source range; a node it kept carries its own, and that is how the two are told apart
  • TypeAwarePass — reads the shared SemanticModel and emits sub-statement text edits keyed by TextRange. it never mutates the AST, because inferred_type binds to the exact parsed node identities

order matters: passes that target the same offset rely on a fixed sequence (e.g. type_is before identity_swap, coalesce before none_chain). the ordered lists live in run_against_source. a complete list of transforms is at crates/by_transforms/src/transforms/mod.rs; each module's /// docs describe the rewrite it performs

stubs

a .byi transpiles to a .pyi, which a checker reads and python never runs. whether a file is a stub is read off the file itself by transpile_typed, since by build hands every source it stages the one config

a pass whose output exists only for running — a runtime check, an entry point, reification, a quoted forward reference — says so through runtime_only, and the driver leaves it out of a stub. a pass that lowers syntax never does: left out, its construct would reach the .pyi as something python cannot parse. the few passes that do both keep the declaration and drop the rest themselves: a stub keeps its imports eager, declares an enum's variants without attaching them, and keeps a mutable default without its guard

splicing

after the passes run, the driver assembles the output in one pass:

  1. hoisted statements (synthesized class defs inserted before the statement that needs them)
  2. a template for each node an AST pass changed. it is printed from the syntax tree with every node below it that kept its source range as a placeholder, and each placeholder passes that node's source through — so a conversion, a quoted forward reference or an extension call a TypeAwarePass lowered inside it is applied there, where printing the statement whole would have dropped it. whether a node changed is read by printing it the same way from the parse and comparing, with the names it spells as the source spelled them: a node whose one change is a name, a parameter a repeated _ numbered, re-emits only that name, and a compound statement whose one change is the statements of a suite re-emits only that suite — its header keeps its source, where the parser stands up what python has no spelling for (a modifier as a decorator, init(...) as a def __init__, a raises clause) and the passes that lower it edit that source. a suite written on the line of its clause (def f(): x) is re-emitted as a block below it, since it has no line of its own to keep. a node a pass builds by parsing text of its own has to forget the ranges into that text, or they would name source they did not come from. between two nodes it passes through, a template keeps the source wherever it prints the same tokens there — the comments, the blank lines, the layout of a signature, and the edits other passes made in that text — and elsewhere keeps the comments ahead of the first token it prints differently and after the last, where it breaks the line too. a comment between two tokens it prints differently has no line of its own in the output, and is not kept. what a template prints is indented by the step its node's own source indents its blocks by, since a file may indent one block by two spaces and another by four
  3. sub-statement text edits, applied with ruff-style first-wins overlap skip — a wider edit wins over a narrower one nested inside it, and a template applies the edits inside the source it passes through. the splice records every edit it writes out, and one it leaves out has been lost — whatever the wider edit prints there stands in for that source, however it spells it — and is a transpile error, unless the wider edit accounts for it: it deletes what it covers, it writes the same thing at the same span, one pass wrote both, or its pass declares that it writes the lost edit's lowering itself (TypeAwarePass::subsumes, a list of Lowerings). a changed node accounts for nothing, and a lowering no pass declares is refused wherever another edit covers it, so a new lowering that lands inside a construct some pass writes itself fails the transpile until that pass takes it on
  4. required_imports prepended (deduped, from-imports merged)
  5. __all__ epilogue appended for export/public modifiers

source maps

transpile_typed_with_map returns a line table (output line → .by line, None for generated lines), composed from the phase-0 table plus the count of generated preamble lines. it powers by run's traceback rewriting and the .by source span on transpiler-error diagnostics. the table is line-level only; the byte-accurate, bidirectional design is in sourcemaps

reverse transforms

basedpython also supports reverse transpilation — converting standard Python back into basedpython syntax. see reverse transforms for details