Skip to content

[interop] Cache actual class and base offset lookups - #115

Open
guitargeek wants to merge 1 commit into
compiler-research:mainfrom
guitargeek:patch-1
Open

guitargeek wants to merge 1 commit into
compiler-research:mainfrom
guitargeek:patch-1

Conversation

@guitargeek

Copy link
Copy Markdown
Collaborator

Every polymorphic object handed to Python goes through GetActualClass(), which demangles the dynamic type's name and resolves it with a full interpreter name lookup, and then through GetBaseOffset(), which walks the inheritance paths. Neither result was cached, so this cost was paid again for every returned object.

Memoize both in the interop wrapper, where the calls are already serialized by the interop lock:

  • GetActualClass per (static class, dynamic type_info). Only outcomes that cannot change are cached: a failed lookup may succeed once more declarations are available, so it is retried.
  • GetBaseOffset per (derived, base) class pair. The object address was already ignored, so the offset is fixed per pair once both records are complete. An incomplete record has no walkable bases (the lookup would read an empty base path), so -- like the error case -- it is treated silently and retried next time rather than frozen into the cache.

For example, getting objects out of a container of base pointers, with auto-downcast to the actual derived class on each access:

import time
import cppjit

cppjit.cppdef(r"""
#include <vector>
struct PerfBase {
    virtual ~PerfBase() = default;
    virtual int value() const { return 1; }
};
struct PerfDerived : PerfBase {
    int value() const override { return 2; }
};
std::vector<PerfBase*> make_perf_vec(int n) {
    static PerfDerived d[64];
    std::vector<PerfBase*> v;
    for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
    return v;
}
""")

vec = cppjit.gbl.make_perf_vec(100)
vec.at(0)  # warm-up
start = time.perf_counter()
for _ in range(2000):
    for i in range(100):
        vec.at(i)  # returns PerfBase*, auto-downcast to PerfDerived
print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")

This takes 0.18 s instead of 0.40 s per 200k accesses (Release build against the pinned CppInterOp, min of 7 repetitions, interleaved A/B against the pristine library, ~2.3x). An instrumented build confirms one GetActualClass() and one GetBaseOffset() call per element access before this change.

If the name lookup is expensive -- the actual class is templated with a non-type template argument, e.g. Impl<float, (EOp)0>, whose spelled name needs interpreter-assisted resolution -- the difference is much larger: a single returned base pointer cost ~420 us and now costs ~0.6 us.

Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.

Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:

- GetActualClass per (static class, dynamic type_info). Only outcomes
  that cannot change are cached: a failed lookup may succeed once more
  declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
  already ignored, so the offset is fixed per pair once both records are
  complete. An incomplete record has no walkable bases (the lookup would
  read an empty base path), so -- like the error case -- it is treated
  silently and retried next time rather than frozen into the cache.

For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:

    import time
    import cppjit

    cppjit.cppdef(r"""
    #include <vector>
    struct PerfBase {
        virtual ~PerfBase() = default;
        virtual int value() const { return 1; }
    };
    struct PerfDerived : PerfBase {
        int value() const override { return 2; }
    };
    std::vector<PerfBase*> make_perf_vec(int n) {
        static PerfDerived d[64];
        std::vector<PerfBase*> v;
        for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
        return v;
    }
    """)

    vec = cppjit.gbl.make_perf_vec(100)
    vec.at(0)  # warm-up
    start = time.perf_counter()
    for _ in range(2000):
        for i in range(100):
            vec.at(i)  # returns PerfBase*, auto-downcast to PerfDerived
    print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")

This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.

If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.
@guitargeek guitargeek self-assigned this Sep 27, 2026
aaronj0 pushed a commit to aaronj0/root that referenced this pull request Oct 1, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.

Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:

- GetActualClass per (static class, dynamic type_info). Only outcomes
  that cannot change are cached: a failed lookup may succeed once more
  declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
  already ignored, so the offset is fixed per pair once both records are
  complete. An incomplete record has no walkable bases (the lookup would
  read an empty base path), so -- like the error case -- it is treated
  silently and retried next time rather than frozen into the cache.

For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:

    import time
    import cppjit

    cppjit.cppdef(r"""
    #include <vector>
    struct PerfBase {
        virtual ~PerfBase() = default;
        virtual int value() const { return 1; }
    };
    struct PerfDerived : PerfBase {
        int value() const override { return 2; }
    };
    std::vector<PerfBase*> make_perf_vec(int n) {
        static PerfDerived d[64];
        std::vector<PerfBase*> v;
        for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
        return v;
    }
    """)

    vec = cppjit.gbl.make_perf_vec(100)
    vec.at(0)  # warm-up
    start = time.perf_counter()
    for _ in range(2000):
        for i in range(100):
            vec.at(i)  # returns PerfBase*, auto-downcast to PerfDerived
    print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")

This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.

If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.

Matches compiler-research/cppjit#115.
aaronj0 pushed a commit to aaronj0/root that referenced this pull request Oct 1, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.

Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:

- GetActualClass per (static class, dynamic type_info). Only outcomes
  that cannot change are cached: a failed lookup may succeed once more
  declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
  already ignored, so the offset is fixed per pair once both records are
  complete. An incomplete record has no walkable bases (the lookup would
  read an empty base path), so -- like the error case -- it is treated
  silently and retried next time rather than frozen into the cache.

For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:

    import time
    import cppjit

    cppjit.cppdef(r"""
    #include <vector>
    struct PerfBase {
        virtual ~PerfBase() = default;
        virtual int value() const { return 1; }
    };
    struct PerfDerived : PerfBase {
        int value() const override { return 2; }
    };
    std::vector<PerfBase*> make_perf_vec(int n) {
        static PerfDerived d[64];
        std::vector<PerfBase*> v;
        for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
        return v;
    }
    """)

    vec = cppjit.gbl.make_perf_vec(100)
    vec.at(0)  # warm-up
    start = time.perf_counter()
    for _ in range(2000):
        for i in range(100):
            vec.at(i)  # returns PerfBase*, auto-downcast to PerfDerived
    print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")

This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.

If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.

Matches compiler-research/cppjit#115.
aaronj0 pushed a commit to aaronj0/root that referenced this pull request Oct 1, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.

Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:

- GetActualClass per (static class, dynamic type_info). Only outcomes
  that cannot change are cached: a failed lookup may succeed once more
  declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
  already ignored, so the offset is fixed per pair once both records are
  complete. An incomplete record has no walkable bases (the lookup would
  read an empty base path), so -- like the error case -- it is treated
  silently and retried next time rather than frozen into the cache.

For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:

    import time
    import cppjit

    cppjit.cppdef(r"""
    #include <vector>
    struct PerfBase {
        virtual ~PerfBase() = default;
        virtual int value() const { return 1; }
    };
    struct PerfDerived : PerfBase {
        int value() const override { return 2; }
    };
    std::vector<PerfBase*> make_perf_vec(int n) {
        static PerfDerived d[64];
        std::vector<PerfBase*> v;
        for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
        return v;
    }
    """)

    vec = cppjit.gbl.make_perf_vec(100)
    vec.at(0)  # warm-up
    start = time.perf_counter()
    for _ in range(2000):
        for i in range(100):
            vec.at(i)  # returns PerfBase*, auto-downcast to PerfDerived
    print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")

This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.

If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.

Matches compiler-research/cppjit#115.
aaronj0 pushed a commit to aaronj0/root that referenced this pull request Oct 2, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.

Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:

- GetActualClass per (static class, dynamic type_info). Only outcomes
  that cannot change are cached: a failed lookup may succeed once more
  declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
  already ignored, so the offset is fixed per pair once both records are
  complete. An incomplete record has no walkable bases (the lookup would
  read an empty base path), so -- like the error case -- it is treated
  silently and retried next time rather than frozen into the cache.

For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:

    import time
    import cppjit

    cppjit.cppdef(r"""
    #include <vector>
    struct PerfBase {
        virtual ~PerfBase() = default;
        virtual int value() const { return 1; }
    };
    struct PerfDerived : PerfBase {
        int value() const override { return 2; }
    };
    std::vector<PerfBase*> make_perf_vec(int n) {
        static PerfDerived d[64];
        std::vector<PerfBase*> v;
        for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
        return v;
    }
    """)

    vec = cppjit.gbl.make_perf_vec(100)
    vec.at(0)  # warm-up
    start = time.perf_counter()
    for _ in range(2000):
        for i in range(100):
            vec.at(i)  # returns PerfBase*, auto-downcast to PerfDerived
    print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")

This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.

If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.

Matches compiler-research/cppjit#115.
aaronj0 pushed a commit that referenced this pull request Oct 3, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.

Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:

- GetActualClass per (static class, dynamic type_info). Only outcomes
  that cannot change are cached: a failed lookup may succeed once more
  declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
  already ignored, so the offset is fixed per pair once both records are
  complete. An incomplete record has no walkable bases (the lookup would
  read an empty base path), so -- like the error case -- it is treated
  silently and retried next time rather than frozen into the cache.

For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:

    import time
    import cppjit

    cppjit.cppdef(r"""
    #include <vector>
    struct PerfBase {
        virtual ~PerfBase() = default;
        virtual int value() const { return 1; }
    };
    struct PerfDerived : PerfBase {
        int value() const override { return 2; }
    };
    std::vector<PerfBase*> make_perf_vec(int n) {
        static PerfDerived d[64];
        std::vector<PerfBase*> v;
        for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
        return v;
    }
    """)

    vec = cppjit.gbl.make_perf_vec(100)
    vec.at(0)  # warm-up
    start = time.perf_counter()
    for _ in range(2000):
        for i in range(100):
            vec.at(i)  # returns PerfBase*, auto-downcast to PerfDerived
    print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")

This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.

If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.

Matches #115.
aaronj0 pushed a commit to root-project/root that referenced this pull request Oct 3, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.

Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:

- GetActualClass per (static class, dynamic type_info). Only outcomes
  that cannot change are cached: a failed lookup may succeed once more
  declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
  already ignored, so the offset is fixed per pair once both records are
  complete. An incomplete record has no walkable bases (the lookup would
  read an empty base path), so -- like the error case -- it is treated
  silently and retried next time rather than frozen into the cache.

For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:

    import time
    import cppjit

    cppjit.cppdef(r"""
    #include <vector>
    struct PerfBase {
        virtual ~PerfBase() = default;
        virtual int value() const { return 1; }
    };
    struct PerfDerived : PerfBase {
        int value() const override { return 2; }
    };
    std::vector<PerfBase*> make_perf_vec(int n) {
        static PerfDerived d[64];
        std::vector<PerfBase*> v;
        for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
        return v;
    }
    """)

    vec = cppjit.gbl.make_perf_vec(100)
    vec.at(0)  # warm-up
    start = time.perf_counter()
    for _ in range(2000):
        for i in range(100):
            vec.at(i)  # returns PerfBase*, auto-downcast to PerfDerived
    print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")

This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.

If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.

Matches compiler-research/cppjit#115.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant