Repository navigation
Support loadSubset deduplication in query collection #836
Description
Activity
I don't think this is an easy thing to solve, especially if you look at the whole spectrum of possible predicates / subsets. When I query for ids / keys, it becomes pretty easy (because we can directly map an id to a record in our collection), whereas more exotic predicates that map to properties can't actually be deduped reliably (since it would be naive to take any cached content as the full truth ... imagine querying for all users that have been active within the last 5 minutes via e.g.
gte(user.lastSeenUnix, Math.floor(Date.now()/1000) - 1000*60*5), by the time I mount a new query with the same predicate, I might have a completely different set of users returned / none returned). You always risk showcasing stale data when you deduplicate, regardless of whatstaleTimeis specified.
I don't think any amount of deduplication work or brainstorming is going to solve the staleness issue we will face with deduplicating ambitious queries (e.g. queries not tied to an identifying property like the key / id).
In case of strict predicates, e.g. querying for ids / keys (e.g.
eq(user.id, "123")orinArray(user.id, ["123", "456"]), it becomes much easier to handle since you already know what you're expecting (user with id 123 or 456).If you're just looking at it from a "grab entity by key / id" angle, it could be pretty easy to solve even on the developer (people using this framework) side.
My current implementation wraps every entity in a makeshift
DbSyncMetaRecord, which resolves to this:type DbSyncMetaRecord<T = unknown> = { id: string | number; record: T; lastUpdate: number; };
The idea is, that every entity / record is added to the collection wrapped in this structure, where lastUpdate is the current unix timestamp from when the data was pulled in by the queryFn from the remote endpoint / API. That way I know exactly when each e.g. user object hit my cache.
Now when I go ahead and mount a second query, the queryFn could find all cache hits via the id filter, compare the lastUpdate timestamp against the current unix and factor into the staleTime and determine, whether the filter should include that specific id or not. If userid
123was mounted 2 minutes ago, but my collection staleTime is 5 minutes, it could be pruned from the predicate that is sent to the API while instantly returning the cached user record (and pulling in the remaining ones async). If it's considered stale (NOW - lastUpdate >= 5 min), it's kept in the predicate so the existing record gets pulled in fresh from the API and is updated with the remaining batch.This is acceptable behavior for my app currently, seeing as I only have a need for querying by id at the moment.
more exotic predicates that map to properties can't actually be deduped reliably (since it would be naive to take any cached content as the full truth ... imagine querying for all users that have been active within the last 5 minutes via e.g. gte(user.lastSeenUnix, Math.floor(Date.now()/1000) - 1000605), by the time I mount a new query with the same predicate, I might have a completely different set of users returned / none returned). You always risk showcasing stale data when you deduplicate, regardless of what staleTime is specified.
@TimFL A query always runs against the local collection. Whether the data in that collection is stale and how often it is "synced" with the actual backend depends on the collection implementation. For example, Electric collections are usually in-sync because the Electric backend pushes updates in real time to the client (when online ofcourse). The query collection is pull-based and thus can contain stale data for a longer period of time, depending on how the tanstack/queries are configured. So it is perfectly normal that a query returns stale data if the data in the collection is stale.
NB: queries are evaluated once, so if your query is
gte(user.lastSeenUnix, Math.floor(Date.now()/1000) - 1000*60*5)thenDate.now()is evaluated once and will not change during the lifetime of your query. It would be interesting though to introduce aNOW()operator in the future which would effectively mean that the predicate changes throughout the query's lifetime. But that's not planned for near future and we shouldn't block the deduplication work based on this as we don't currently support it.The main challenge with the deduplication of queries is to tie the lifetime of the deduper's internal state to the lifetime of the queries that load/reference that data. I.e. detect when data is no longer referenced by any queries in order to GC it. And make sure that if the original query that loaded some data (and is still "syncing" that data) is GCed that the deduplicated queries still receive updates (i.e one of the remaining queries needs to be promoted to be the one that "syncs" that data from the backend).
@kevin-dp could the deduping happen at the api layer? Still create the query observer, but before calling the api we utilize
loadSubsetto dictate the api call? AnunloadSubsetwould need to be added, that would get called on query unsubscribe. The deduplicated load subset would need to keep track of references per predicate but the flow could be something like:Flow:
- Query subscribes -> creates its own independent
QueryObserver(unchanged) - Observer's queryFn fires -> passes predicate to
DeduplicatedLoadSubsetbefore hitting the API
- Already loaded by another query? -> resolve from local collection, skip API call
- Partially overlapping? ->minusWherePredicatescomputes the diff, only fetches what's missing
- In-flight request covers it? -> waits for that instead of duplicating
- No coverage? -> normal API call - Query unsubscribes -> row cleanup via existing
queryToRows/rowToQueries+DeduplicatedLoadSubsetdecrements predicate reference count - Superset query GC'd? -> subset's data stays (owns its own rows), next refetch finds no coverage _> real API call. No stale data.
- Query subscribes -> creates its own independent
Hi @goatrenterguy,
I believe steps 1 and 2 are essentially how it used to be implemented and that led to the problems described in this issue. Steps 3 and 4 indeed seem similar to what i had in mind, namely explicitly tracking reference counts, except you're tracking it on a per row basis and we were thinking of tracking it on a per-query basis.
Closing as fixed on current main. Model-backed oracles now cover identical in-flight demand sharing, final exact acquisition release, overlapping-owner row retention, and startup reference-count histories. Thanks @kevin-dp for the lifetime-ownership framing; the current implementation follows that durable invariant.
In #835 we disabled the deduplication of
loadSubsetcalls inside the query collection because there were some problems with garbage collection of queries which would lead to deduplicated queries to use stale data from the local collection (because the query they depend on was the one loading updates from the backend but is now GCed).However, we do need to optimize the query collection such that we do not fetch data we already have locally and integrate it smartly with GC such that we detect when data becomes stale and needs to be fetched from the backend. The main problem came from the fact that we did reference counting on a per-row basis, we're thinking that it would be more maintainable on a per-query basis, i.e. if query 1 loads some data and query 2 is deduplicated because of query 1, then query 1's ref count is essentially 2 (meaning that it can only be GCed after query 2 has been GCed).