[interop] Cache actual class and base offset lookups - #115
Open
guitargeek wants to merge 1 commit into
Open
guitargeek wants to merge 1 commit into
guitargeek wants to merge 1 commit into
Conversation
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.
Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:
- GetActualClass per (static class, dynamic type_info). Only outcomes
that cannot change are cached: a failed lookup may succeed once more
declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
already ignored, so the offset is fixed per pair once both records are
complete. An incomplete record has no walkable bases (the lookup would
read an empty base path), so -- like the error case -- it is treated
silently and retried next time rather than frozen into the cache.
For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:
import time
import cppjit
cppjit.cppdef(r"""
#include <vector>
struct PerfBase {
virtual ~PerfBase() = default;
virtual int value() const { return 1; }
};
struct PerfDerived : PerfBase {
int value() const override { return 2; }
};
std::vector<PerfBase*> make_perf_vec(int n) {
static PerfDerived d[64];
std::vector<PerfBase*> v;
for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
return v;
}
""")
vec = cppjit.gbl.make_perf_vec(100)
vec.at(0) # warm-up
start = time.perf_counter()
for _ in range(2000):
for i in range(100):
vec.at(i) # returns PerfBase*, auto-downcast to PerfDerived
print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")
This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.
If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.
aaronj0
pushed a commit
to aaronj0/root
that referenced
this pull request
Oct 1, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.
Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:
- GetActualClass per (static class, dynamic type_info). Only outcomes
that cannot change are cached: a failed lookup may succeed once more
declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
already ignored, so the offset is fixed per pair once both records are
complete. An incomplete record has no walkable bases (the lookup would
read an empty base path), so -- like the error case -- it is treated
silently and retried next time rather than frozen into the cache.
For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:
import time
import cppjit
cppjit.cppdef(r"""
#include <vector>
struct PerfBase {
virtual ~PerfBase() = default;
virtual int value() const { return 1; }
};
struct PerfDerived : PerfBase {
int value() const override { return 2; }
};
std::vector<PerfBase*> make_perf_vec(int n) {
static PerfDerived d[64];
std::vector<PerfBase*> v;
for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
return v;
}
""")
vec = cppjit.gbl.make_perf_vec(100)
vec.at(0) # warm-up
start = time.perf_counter()
for _ in range(2000):
for i in range(100):
vec.at(i) # returns PerfBase*, auto-downcast to PerfDerived
print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")
This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.
If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.
Matches compiler-research/cppjit#115.
aaronj0
pushed a commit
to aaronj0/root
that referenced
this pull request
Oct 1, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.
Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:
- GetActualClass per (static class, dynamic type_info). Only outcomes
that cannot change are cached: a failed lookup may succeed once more
declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
already ignored, so the offset is fixed per pair once both records are
complete. An incomplete record has no walkable bases (the lookup would
read an empty base path), so -- like the error case -- it is treated
silently and retried next time rather than frozen into the cache.
For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:
import time
import cppjit
cppjit.cppdef(r"""
#include <vector>
struct PerfBase {
virtual ~PerfBase() = default;
virtual int value() const { return 1; }
};
struct PerfDerived : PerfBase {
int value() const override { return 2; }
};
std::vector<PerfBase*> make_perf_vec(int n) {
static PerfDerived d[64];
std::vector<PerfBase*> v;
for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
return v;
}
""")
vec = cppjit.gbl.make_perf_vec(100)
vec.at(0) # warm-up
start = time.perf_counter()
for _ in range(2000):
for i in range(100):
vec.at(i) # returns PerfBase*, auto-downcast to PerfDerived
print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")
This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.
If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.
Matches compiler-research/cppjit#115.
aaronj0
pushed a commit
to aaronj0/root
that referenced
this pull request
Oct 1, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.
Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:
- GetActualClass per (static class, dynamic type_info). Only outcomes
that cannot change are cached: a failed lookup may succeed once more
declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
already ignored, so the offset is fixed per pair once both records are
complete. An incomplete record has no walkable bases (the lookup would
read an empty base path), so -- like the error case -- it is treated
silently and retried next time rather than frozen into the cache.
For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:
import time
import cppjit
cppjit.cppdef(r"""
#include <vector>
struct PerfBase {
virtual ~PerfBase() = default;
virtual int value() const { return 1; }
};
struct PerfDerived : PerfBase {
int value() const override { return 2; }
};
std::vector<PerfBase*> make_perf_vec(int n) {
static PerfDerived d[64];
std::vector<PerfBase*> v;
for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
return v;
}
""")
vec = cppjit.gbl.make_perf_vec(100)
vec.at(0) # warm-up
start = time.perf_counter()
for _ in range(2000):
for i in range(100):
vec.at(i) # returns PerfBase*, auto-downcast to PerfDerived
print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")
This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.
If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.
Matches compiler-research/cppjit#115.
aaronj0
pushed a commit
to aaronj0/root
that referenced
this pull request
Oct 2, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.
Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:
- GetActualClass per (static class, dynamic type_info). Only outcomes
that cannot change are cached: a failed lookup may succeed once more
declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
already ignored, so the offset is fixed per pair once both records are
complete. An incomplete record has no walkable bases (the lookup would
read an empty base path), so -- like the error case -- it is treated
silently and retried next time rather than frozen into the cache.
For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:
import time
import cppjit
cppjit.cppdef(r"""
#include <vector>
struct PerfBase {
virtual ~PerfBase() = default;
virtual int value() const { return 1; }
};
struct PerfDerived : PerfBase {
int value() const override { return 2; }
};
std::vector<PerfBase*> make_perf_vec(int n) {
static PerfDerived d[64];
std::vector<PerfBase*> v;
for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
return v;
}
""")
vec = cppjit.gbl.make_perf_vec(100)
vec.at(0) # warm-up
start = time.perf_counter()
for _ in range(2000):
for i in range(100):
vec.at(i) # returns PerfBase*, auto-downcast to PerfDerived
print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")
This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.
If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.
Matches compiler-research/cppjit#115.
aaronj0
pushed a commit
that referenced
this pull request
Oct 3, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.
Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:
- GetActualClass per (static class, dynamic type_info). Only outcomes
that cannot change are cached: a failed lookup may succeed once more
declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
already ignored, so the offset is fixed per pair once both records are
complete. An incomplete record has no walkable bases (the lookup would
read an empty base path), so -- like the error case -- it is treated
silently and retried next time rather than frozen into the cache.
For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:
import time
import cppjit
cppjit.cppdef(r"""
#include <vector>
struct PerfBase {
virtual ~PerfBase() = default;
virtual int value() const { return 1; }
};
struct PerfDerived : PerfBase {
int value() const override { return 2; }
};
std::vector<PerfBase*> make_perf_vec(int n) {
static PerfDerived d[64];
std::vector<PerfBase*> v;
for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
return v;
}
""")
vec = cppjit.gbl.make_perf_vec(100)
vec.at(0) # warm-up
start = time.perf_counter()
for _ in range(2000):
for i in range(100):
vec.at(i) # returns PerfBase*, auto-downcast to PerfDerived
print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")
This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.
If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.
Matches #115.
aaronj0
pushed a commit
to root-project/root
that referenced
this pull request
Oct 3, 2026
Every polymorphic object handed to Python goes through GetActualClass(),
which demangles the dynamic type's name and resolves it with a full
interpreter name lookup, and then through GetBaseOffset(), which walks
the inheritance paths. Neither result was cached, so this cost was paid
again for every returned object.
Memoize both in the interop wrapper, where the calls are already
serialized by the interop lock:
- GetActualClass per (static class, dynamic type_info). Only outcomes
that cannot change are cached: a failed lookup may succeed once more
declarations are available, so it is retried.
- GetBaseOffset per (derived, base) class pair. The object address was
already ignored, so the offset is fixed per pair once both records are
complete. An incomplete record has no walkable bases (the lookup would
read an empty base path), so -- like the error case -- it is treated
silently and retried next time rather than frozen into the cache.
For example, getting objects out of a container of base pointers, with
auto-downcast to the actual derived class on each access:
import time
import cppjit
cppjit.cppdef(r"""
#include <vector>
struct PerfBase {
virtual ~PerfBase() = default;
virtual int value() const { return 1; }
};
struct PerfDerived : PerfBase {
int value() const override { return 2; }
};
std::vector<PerfBase*> make_perf_vec(int n) {
static PerfDerived d[64];
std::vector<PerfBase*> v;
for (int i = 0; i < n; ++i) v.push_back(&d[i % 64]);
return v;
}
""")
vec = cppjit.gbl.make_perf_vec(100)
vec.at(0) # warm-up
start = time.perf_counter()
for _ in range(2000):
for i in range(100):
vec.at(i) # returns PerfBase*, auto-downcast to PerfDerived
print(f"200k vector<PerfBase*>.at() calls: {time.perf_counter() - start:.2f} s")
This takes 0.18 s instead of 0.40 s per 200k accesses (Release build
against the pinned CppInterOp, min of 7 repetitions, interleaved A/B
against the pristine library, ~2.3x). An instrumented build confirms one
GetActualClass() and one GetBaseOffset() call per element access before
this change.
If the name lookup is expensive -- the actual class is templated with a
non-type template argument, e.g. `Impl<float, (EOp)0>`, whose spelled
name needs interpreter-assisted resolution -- the difference is much
larger: a single returned base pointer cost ~420 us and now costs
~0.6 us.
Matches compiler-research/cppjit#115.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every polymorphic object handed to Python goes through GetActualClass(), which demangles the dynamic type's name and resolves it with a full interpreter name lookup, and then through GetBaseOffset(), which walks the inheritance paths. Neither result was cached, so this cost was paid again for every returned object.
Memoize both in the interop wrapper, where the calls are already serialized by the interop lock:
For example, getting objects out of a container of base pointers, with auto-downcast to the actual derived class on each access:
This takes 0.18 s instead of 0.40 s per 200k accesses (Release build against the pinned CppInterOp, min of 7 repetitions, interleaved A/B against the pristine library, ~2.3x). An instrumented build confirms one GetActualClass() and one GetBaseOffset() call per element access before this change.
If the name lookup is expensive -- the actual class is templated with a non-type template argument, e.g.
Impl<float, (EOp)0>, whose spelled name needs interpreter-assisted resolution -- the difference is much larger: a single returned base pointer cost ~420 us and now costs ~0.6 us.