Skip to content

Support loadSubset deduplication in query collection #836

Description

@kevin-dp

In #835 we disabled the deduplication of loadSubset calls inside the query collection because there were some problems with garbage collection of queries which would lead to deduplicated queries to use stale data from the local collection (because the query they depend on was the one loading updates from the backend but is now GCed).

However, we do need to optimize the query collection such that we do not fetch data we already have locally and integrate it smartly with GC such that we detect when data becomes stale and needs to be fetched from the backend. The main problem came from the fact that we did reference counting on a per-row basis, we're thinking that it would be more maintainable on a per-query basis, i.e. if query 1 loads some data and query 2 is deduplicated because of query 1, then query 1's ref count is essentially 2 (meaning that it can only be GCed after query 2 has been GCed).

Activity

  1. TimFL commented on Nov 17, 2025

    @TimFL
    Contributor

    I don't think this is an easy thing to solve, especially if you look at the whole spectrum of possible predicates / subsets. When I query for ids / keys, it becomes pretty easy (because we can directly map an id to a record in our collection), whereas more exotic predicates that map to properties can't actually be deduped reliably (since it would be naive to take any cached content as the full truth ... imagine querying for all users that have been active within the last 5 minutes via e.g. gte(user.lastSeenUnix, Math.floor(Date.now()/1000) - 1000*60*5), by the time I mount a new query with the same predicate, I might have a completely different set of users returned / none returned). You always risk showcasing stale data when you deduplicate, regardless of what staleTime is specified.
    I don't think any amount of deduplication work or brainstorming is going to solve the staleness issue we will face with deduplicating ambitious queries (e.g. queries not tied to an identifying property like the key / id).


    In case of strict predicates, e.g. querying for ids / keys (e.g. eq(user.id, "123") or inArray(user.id, ["123", "456"]), it becomes much easier to handle since you already know what you're expecting (user with id 123 or 456).

    If you're just looking at it from a "grab entity by key / id" angle, it could be pretty easy to solve even on the developer (people using this framework) side.

    My current implementation wraps every entity in a makeshift DbSyncMetaRecord, which resolves to this:

    type DbSyncMetaRecord<T = unknown> = {
    	id: string | number;
    	record: T;
    	lastUpdate: number;
    };

    The idea is, that every entity / record is added to the collection wrapped in this structure, where lastUpdate is the current unix timestamp from when the data was pulled in by the queryFn from the remote endpoint / API. That way I know exactly when each e.g. user object hit my cache.

    Now when I go ahead and mount a second query, the queryFn could find all cache hits via the id filter, compare the lastUpdate timestamp against the current unix and factor into the staleTime and determine, whether the filter should include that specific id or not. If userid 123 was mounted 2 minutes ago, but my collection staleTime is 5 minutes, it could be pruned from the predicate that is sent to the API while instantly returning the cached user record (and pulling in the remaining ones async). If it's considered stale (NOW - lastUpdate >= 5 min), it's kept in the predicate so the existing record gets pulled in fresh from the API and is updated with the remaining batch.

    This is acceptable behavior for my app currently, seeing as I only have a need for querying by id at the moment.

  2. kevin-dp commented on Nov 24, 2025

    @kevin-dp
    ContributorAuthor

    more exotic predicates that map to properties can't actually be deduped reliably (since it would be naive to take any cached content as the full truth ... imagine querying for all users that have been active within the last 5 minutes via e.g. gte(user.lastSeenUnix, Math.floor(Date.now()/1000) - 1000605), by the time I mount a new query with the same predicate, I might have a completely different set of users returned / none returned). You always risk showcasing stale data when you deduplicate, regardless of what staleTime is specified.

    @TimFL A query always runs against the local collection. Whether the data in that collection is stale and how often it is "synced" with the actual backend depends on the collection implementation. For example, Electric collections are usually in-sync because the Electric backend pushes updates in real time to the client (when online ofcourse). The query collection is pull-based and thus can contain stale data for a longer period of time, depending on how the tanstack/queries are configured. So it is perfectly normal that a query returns stale data if the data in the collection is stale.

    NB: queries are evaluated once, so if your query is gte(user.lastSeenUnix, Math.floor(Date.now()/1000) - 1000*60*5) then Date.now() is evaluated once and will not change during the lifetime of your query. It would be interesting though to introduce a NOW() operator in the future which would effectively mean that the predicate changes throughout the query's lifetime. But that's not planned for near future and we shouldn't block the deduplication work based on this as we don't currently support it.

    The main challenge with the deduplication of queries is to tie the lifetime of the deduper's internal state to the lifetime of the queries that load/reference that data. I.e. detect when data is no longer referenced by any queries in order to GC it. And make sure that if the original query that loaded some data (and is still "syncing" that data) is GCed that the deduplicated queries still receive updates (i.e one of the remaining queries needs to be promoted to be the one that "syncs" that data from the backend).

  3. goatrenterguy commented on Feb 6, 2026

    @goatrenterguy
    Contributor

    @kevin-dp could the deduping happen at the api layer? Still create the query observer, but before calling the api we utilize loadSubset to dictate the api call? An unloadSubset would need to be added, that would get called on query unsubscribe. The deduplicated load subset would need to keep track of references per predicate but the flow could be something like:

    Flow:

    1. Query subscribes -> creates its own independent QueryObserver (unchanged)
    2. Observer's queryFn fires -> passes predicate to DeduplicatedLoadSubset before hitting the API
      - Already loaded by another query? -> resolve from local collection, skip API call
      - Partially overlapping? -> minusWherePredicates computes the diff, only fetches what's missing
      - In-flight request covers it? -> waits for that instead of duplicating
      - No coverage? -> normal API call
    3. Query unsubscribes -> row cleanup via existing queryToRows/rowToQueries + DeduplicatedLoadSubset decrements predicate reference count
    4. Superset query GC'd? -> subset's data stays (owns its own rows), next refetch finds no coverage _> real API call. No stale data.
  4. kevin-dp commented on Feb 9, 2026

    @kevin-dp
    ContributorAuthor

    Hi @goatrenterguy,

    I believe steps 1 and 2 are essentially how it used to be implemented and that led to the problems described in this issue. Steps 3 and 4 indeed seem similar to what i had in mind, namely explicitly tracking reference counts, except you're tracking it on a per row basis and we were thinking of tracking it on a per-query basis.

  5. KyleAMathews commented on Sep 16, 2026

    @KyleAMathews
    Collaborator

    Closing as fixed on current main. Model-backed oracles now cover identical in-flight demand sharing, final exact acquisition release, overlapping-owner row retention, and startup reference-count histories. Thanks @kevin-dp for the lifetime-ownership framing; the current implementation follows that durable invariant.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions