Skip to content

[interop] Cache types resolved through the typeof trampoline - #117

Open
guitargeek wants to merge 1 commit into
compiler-research:mainfrom
guitargeek:patch-3
Open

guitargeek wants to merge 1 commit into
compiler-research:mainfrom
guitargeek:patch-3

Conversation

@guitargeek

Copy link
Copy Markdown
Collaborator

GetType() resolves names that Cpp::GetType() cannot handle by declaring using <unique id> = __typeof__(<name>); in the interpreter. Without a cache, every call parses and compiles code and adds one more alias to the AST, growing the interpreter state unboundedly for repeated names.

Memoize the resolved types by name. A name that fails to resolve may succeed once more declarations are available, so failures are retried.

Upstream motivation: in ROOT, reading a TTree array branch creates the converter for the branch's element pointer type (e.g. "Double_t*") on every access, so an event loop went through the trampoline for every entry. The original reproducer:

import array
import time
import ROOT
n = array.array("i", [3])
e = array.array("d", [1.0, 2.0, 3.0])
tree = ROOT.TTree("tree", "")
tree.Branch("n", n, "n/I")
tree.Branch("e", e, "e[n]/D")
for _ in range(10000):
    tree.Fill()
start = time.perf_counter()
for event in tree:
    event.e  # array branch: the converter for "Double_t*" is created on each access
print(f"loop over 10000 entries reading an array branch: {time.perf_counter() - start:.2f} s")

This takes 0.11 s instead of 1.66 s with this change.

On performance in cppjit proper, the honest result of validating this is that no equivalent hot path exists here: with an instrumented build, the full test suite (559 tests) and targeted scenarios (STL container template instantiation, low-level views, callbacks through std::function, operator dispatch, exception translation) produce exactly one trampoline invocation per process, the one-time resolution of std::exception. Spellings like "double*" are resolved by Cpp::GetType directly. So there is no cppjit-side number to quote; the cache prevents AST alias growth for consumers that do drive this path repeatedly, at the cost of one map lookup per call.

GetType() resolves names that Cpp::GetType() cannot handle by declaring
`using <unique id> = __typeof__(<name>);` in the interpreter. Without a
cache, every call parses and compiles code and adds one more alias to
the AST, growing the interpreter state unboundedly for repeated names.

Memoize the resolved types by name. A name that fails to resolve may
succeed once more declarations are available, so failures are retried.

Upstream motivation: in ROOT, reading a TTree array branch creates the
converter for the branch's element pointer type (e.g. "Double_t*") on
every access, so an event loop went through the trampoline for every
entry. The original reproducer:

    import array
    import time
    import ROOT
    n = array.array("i", [3])
    e = array.array("d", [1.0, 2.0, 3.0])
    tree = ROOT.TTree("tree", "")
    tree.Branch("n", n, "n/I")
    tree.Branch("e", e, "e[n]/D")
    for _ in range(10000):
        tree.Fill()
    start = time.perf_counter()
    for event in tree:
        event.e  # array branch: the converter for "Double_t*" is created on each access
    print(f"loop over 10000 entries reading an array branch: {time.perf_counter() - start:.2f} s")

This takes 0.11 s instead of 1.66 s with this change.

On performance in cppjit proper, the honest result of validating this is
that no equivalent hot path exists here: with an instrumented build, the
full test suite (559 tests) and targeted scenarios (STL container
template instantiation, low-level views, callbacks through
std::function, operator dispatch, exception translation) produce exactly
one trampoline invocation per process, the one-time resolution of
std::exception. Spellings like "double*" are resolved by Cpp::GetType
directly. So there is no cppjit-side number to quote; the cache prevents
AST alias growth for consumers that do drive this path repeatedly, at
the cost of one map lookup per call.
@guitargeek guitargeek self-assigned this Sep 27, 2026
// new alias to the AST, so memoize the resolved types by name. A name that
// fails to resolve may succeed once more declarations are available, so
// failures are retried.
static std::map<std::string, TCppType_t> s_slow_type_cache;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can probably put as a key the const char* pointer to "__cppjit_interop_GetType_" + std::to_string(var_count++); which should exist somewhere in the AST (or even the Decl*).

aaronj0 pushed a commit to aaronj0/root that referenced this pull request Oct 1, 2026
GetType() resolves names that Cpp::GetType() cannot handle by declaring
`using <unique id> = __typeof__(<name>);` in the interpreter. Without a
cache, every call parses and compiles code and adds one more alias to
the AST, growing the interpreter state unboundedly for repeated names.

Memoize the resolved types by name. A name that fails to resolve may
succeed once more declarations are available, so failures are retried.

Upstream motivation: in ROOT, reading a TTree array branch creates the
converter for the branch's element pointer type (e.g. "Double_t*") on
every access, so an event loop went through the trampoline for every
entry. The original reproducer:

    import array
    import time
    import ROOT
    n = array.array("i", [3])
    e = array.array("d", [1.0, 2.0, 3.0])
    tree = ROOT.TTree("tree", "")
    tree.Branch("n", n, "n/I")
    tree.Branch("e", e, "e[n]/D")
    for _ in range(10000):
        tree.Fill()
    start = time.perf_counter()
    for event in tree:
        event.e  # array branch: the converter for "Double_t*" is created on each access
    print(f"loop over 10000 entries reading an array branch: {time.perf_counter() - start:.2f} s")

This takes 0.11 s instead of 1.66 s with this change.

On performance in cppjit proper, the honest result of validating this is
that no equivalent hot path exists here: with an instrumented build, the
full test suite (559 tests) and targeted scenarios (STL container
template instantiation, low-level views, callbacks through
std::function, operator dispatch, exception translation) produce exactly
one trampoline invocation per process, the one-time resolution of
std::exception. Spellings like "double*" are resolved by Cpp::GetType
directly. So there is no cppjit-side number to quote; the cache prevents
AST alias growth for consumers that do drive this path repeatedly, at
the cost of one map lookup per call.

Matches compiler-research/cppjit#117.
aaronj0 pushed a commit to aaronj0/root that referenced this pull request Oct 1, 2026
GetType() resolves names that Cpp::GetType() cannot handle by declaring
`using <unique id> = __typeof__(<name>);` in the interpreter. Without a
cache, every call parses and compiles code and adds one more alias to
the AST, growing the interpreter state unboundedly for repeated names.

Memoize the resolved types by name. A name that fails to resolve may
succeed once more declarations are available, so failures are retried.

Upstream motivation: in ROOT, reading a TTree array branch creates the
converter for the branch's element pointer type (e.g. "Double_t*") on
every access, so an event loop went through the trampoline for every
entry. The original reproducer:

    import array
    import time
    import ROOT
    n = array.array("i", [3])
    e = array.array("d", [1.0, 2.0, 3.0])
    tree = ROOT.TTree("tree", "")
    tree.Branch("n", n, "n/I")
    tree.Branch("e", e, "e[n]/D")
    for _ in range(10000):
        tree.Fill()
    start = time.perf_counter()
    for event in tree:
        event.e  # array branch: the converter for "Double_t*" is created on each access
    print(f"loop over 10000 entries reading an array branch: {time.perf_counter() - start:.2f} s")

This takes 0.11 s instead of 1.66 s with this change.

On performance in cppjit proper, the honest result of validating this is
that no equivalent hot path exists here: with an instrumented build, the
full test suite (559 tests) and targeted scenarios (STL container
template instantiation, low-level views, callbacks through
std::function, operator dispatch, exception translation) produce exactly
one trampoline invocation per process, the one-time resolution of
std::exception. Spellings like "double*" are resolved by Cpp::GetType
directly. So there is no cppjit-side number to quote; the cache prevents
AST alias growth for consumers that do drive this path repeatedly, at
the cost of one map lookup per call.

Matches compiler-research/cppjit#117.
aaronj0 pushed a commit to aaronj0/root that referenced this pull request Oct 1, 2026
GetType() resolves names that Cpp::GetType() cannot handle by declaring
`using <unique id> = __typeof__(<name>);` in the interpreter. Without a
cache, every call parses and compiles code and adds one more alias to
the AST, growing the interpreter state unboundedly for repeated names.

Memoize the resolved types by name. A name that fails to resolve may
succeed once more declarations are available, so failures are retried.

Upstream motivation: in ROOT, reading a TTree array branch creates the
converter for the branch's element pointer type (e.g. "Double_t*") on
every access, so an event loop went through the trampoline for every
entry. The original reproducer:

    import array
    import time
    import ROOT
    n = array.array("i", [3])
    e = array.array("d", [1.0, 2.0, 3.0])
    tree = ROOT.TTree("tree", "")
    tree.Branch("n", n, "n/I")
    tree.Branch("e", e, "e[n]/D")
    for _ in range(10000):
        tree.Fill()
    start = time.perf_counter()
    for event in tree:
        event.e  # array branch: the converter for "Double_t*" is created on each access
    print(f"loop over 10000 entries reading an array branch: {time.perf_counter() - start:.2f} s")

This takes 0.11 s instead of 1.66 s with this change.

On performance in cppjit proper, the honest result of validating this is
that no equivalent hot path exists here: with an instrumented build, the
full test suite (559 tests) and targeted scenarios (STL container
template instantiation, low-level views, callbacks through
std::function, operator dispatch, exception translation) produce exactly
one trampoline invocation per process, the one-time resolution of
std::exception. Spellings like "double*" are resolved by Cpp::GetType
directly. So there is no cppjit-side number to quote; the cache prevents
AST alias growth for consumers that do drive this path repeatedly, at
the cost of one map lookup per call.

Matches compiler-research/cppjit#117.
aaronj0 pushed a commit to aaronj0/root that referenced this pull request Oct 2, 2026
GetType() resolves names that Cpp::GetType() cannot handle by declaring
`using <unique id> = __typeof__(<name>);` in the interpreter. Without a
cache, every call parses and compiles code and adds one more alias to
the AST, growing the interpreter state unboundedly for repeated names.

Memoize the resolved types by name. A name that fails to resolve may
succeed once more declarations are available, so failures are retried.

Upstream motivation: in ROOT, reading a TTree array branch creates the
converter for the branch's element pointer type (e.g. "Double_t*") on
every access, so an event loop went through the trampoline for every
entry. The original reproducer:

    import array
    import time
    import ROOT
    n = array.array("i", [3])
    e = array.array("d", [1.0, 2.0, 3.0])
    tree = ROOT.TTree("tree", "")
    tree.Branch("n", n, "n/I")
    tree.Branch("e", e, "e[n]/D")
    for _ in range(10000):
        tree.Fill()
    start = time.perf_counter()
    for event in tree:
        event.e  # array branch: the converter for "Double_t*" is created on each access
    print(f"loop over 10000 entries reading an array branch: {time.perf_counter() - start:.2f} s")

This takes 0.11 s instead of 1.66 s with this change.

On performance in cppjit proper, the honest result of validating this is
that no equivalent hot path exists here: with an instrumented build, the
full test suite (559 tests) and targeted scenarios (STL container
template instantiation, low-level views, callbacks through
std::function, operator dispatch, exception translation) produce exactly
one trampoline invocation per process, the one-time resolution of
std::exception. Spellings like "double*" are resolved by Cpp::GetType
directly. So there is no cppjit-side number to quote; the cache prevents
AST alias growth for consumers that do drive this path repeatedly, at
the cost of one map lookup per call.

Matches compiler-research/cppjit#117.
aaronj0 pushed a commit that referenced this pull request Oct 3, 2026
GetType() resolves names that Cpp::GetType() cannot handle by declaring
`using <unique id> = __typeof__(<name>);` in the interpreter. Without a
cache, every call parses and compiles code and adds one more alias to
the AST, growing the interpreter state unboundedly for repeated names.

Memoize the resolved types by name. A name that fails to resolve may
succeed once more declarations are available, so failures are retried.

Upstream motivation: in ROOT, reading a TTree array branch creates the
converter for the branch's element pointer type (e.g. "Double_t*") on
every access, so an event loop went through the trampoline for every
entry. The original reproducer:

    import array
    import time
    import ROOT
    n = array.array("i", [3])
    e = array.array("d", [1.0, 2.0, 3.0])
    tree = ROOT.TTree("tree", "")
    tree.Branch("n", n, "n/I")
    tree.Branch("e", e, "e[n]/D")
    for _ in range(10000):
        tree.Fill()
    start = time.perf_counter()
    for event in tree:
        event.e  # array branch: the converter for "Double_t*" is created on each access
    print(f"loop over 10000 entries reading an array branch: {time.perf_counter() - start:.2f} s")

This takes 0.11 s instead of 1.66 s with this change.

On performance in cppjit proper, the honest result of validating this is
that no equivalent hot path exists here: with an instrumented build, the
full test suite (559 tests) and targeted scenarios (STL container
template instantiation, low-level views, callbacks through
std::function, operator dispatch, exception translation) produce exactly
one trampoline invocation per process, the one-time resolution of
std::exception. Spellings like "double*" are resolved by Cpp::GetType
directly. So there is no cppjit-side number to quote; the cache prevents
AST alias growth for consumers that do drive this path repeatedly, at
the cost of one map lookup per call.

Matches #117.
aaronj0 pushed a commit to root-project/root that referenced this pull request Oct 3, 2026
GetType() resolves names that Cpp::GetType() cannot handle by declaring
`using <unique id> = __typeof__(<name>);` in the interpreter. Without a
cache, every call parses and compiles code and adds one more alias to
the AST, growing the interpreter state unboundedly for repeated names.

Memoize the resolved types by name. A name that fails to resolve may
succeed once more declarations are available, so failures are retried.

Upstream motivation: in ROOT, reading a TTree array branch creates the
converter for the branch's element pointer type (e.g. "Double_t*") on
every access, so an event loop went through the trampoline for every
entry. The original reproducer:

    import array
    import time
    import ROOT
    n = array.array("i", [3])
    e = array.array("d", [1.0, 2.0, 3.0])
    tree = ROOT.TTree("tree", "")
    tree.Branch("n", n, "n/I")
    tree.Branch("e", e, "e[n]/D")
    for _ in range(10000):
        tree.Fill()
    start = time.perf_counter()
    for event in tree:
        event.e  # array branch: the converter for "Double_t*" is created on each access
    print(f"loop over 10000 entries reading an array branch: {time.perf_counter() - start:.2f} s")

This takes 0.11 s instead of 1.66 s with this change.

On performance in cppjit proper, the honest result of validating this is
that no equivalent hot path exists here: with an instrumented build, the
full test suite (559 tests) and targeted scenarios (STL container
template instantiation, low-level views, callbacks through
std::function, operator dispatch, exception translation) produce exactly
one trampoline invocation per process, the one-time resolution of
std::exception. Spellings like "double*" are resolved by Cpp::GetType
directly. So there is no cppjit-side number to quote; the cache prevents
AST alias growth for consumers that do drive this path repeatedly, at
the cost of one map lookup per call.

Matches compiler-research/cppjit#117.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants