Skip to content

Shared design for DataTree tables with duckdb-zarr: one schema per node #262

Description

@alxmrs

duckdb-zarr is adding support for Zarr stores with nested groups, such as those written by xarray.DataTree.to_zarr(). We'd like both projects to turn the same store into the same tables, with the same names, so this issue proposes a shared design. It builds on #82.

The duckdb-zarr version is decision 8 in xqlsystems/duckdb-zarr#52 (design only, no code yet).

Proposal

  1. A table is one dimension group in one node. It's identified by (node path, dims). Arrays in different nodes never share a table, even when their dimension names match. This is the DataTree model: simulation/coarse and simulation/fine can both have foo(x, y) with different lengths of x.
  2. Tables inside a node are named with the existing rules. default_table_name ("_".join(dims) in dimension order, scalar for none) and resolve_table_names (overrides keyed by the dims tuple, case-insensitive collision check), applied per node.
  3. A node maps to a schema. Tree → catalog, node path → schema, dimension group → table:
    • The root node is the default schema: main in DuckDB, public in DataFusion.
    • A nested path is one quoted schema name: sim."simulation/fine".x_y.
    • In DataFusion this is a CatalogProvider whose SchemaProviders are the nodes, as suggested in Support for xarray-datatrees? #82.
  4. Coordinates are inherited like xarray. A node's tables include coordinates from its ancestors for dimensions the node shares with them, the same view dt["simulation/fine"].to_dataset() gives.
  5. Opening one node uses xarray's word. duckdb-zarr adds read_zarr(store, group := 'simulation/fine'), matching xr.open_zarr(store, group=...). With no group, only the root node is read.

Questions for xarray-sql

  • Does a catalog per tree and a schema per node fit the planned from_datatree API? Or would you rather have flat table names in one schema (for example simulation_fine__x_y)?
  • How should the root node look? Is it the default schema, or a schema named after the tree?
  • Should table_names overrides for a tree be keyed by (node_path, dims), with the per-node dims keys kept as they are?
  • Is anything in xarray-sql's pivot tied to one flat Dataset that would make inherited coordinates hard?

Why now

The immediate user is AnnData (xqlsystems/duckdb-zarr#40). An AnnData store is a tree: obs and var are nodes, and X, layers and obsm share their axes. The AnnData-specific parts come after this: naming its axes and decoding its categorical, nullable and sparse groups. Agreeing on the tree mapping first means both projects expose AnnData the same way.

Activity

  1. alxmrs commented on Sep 29, 2026

    @alxmrs
    MemberAuthor

    Update from the duckdb-zarr side: duckdb-zarr won't mount a store as a catalog itself, at least for now.

    We spiked a CALL zarr_attach(store, alias) that creates one schema per node and one view per dimension group. The DuckDB 1.5.5 C API has no storage-extension hook and can't run SQL on the calling connection, so the procedure needs a second connection that it opens at load and keeps. That connection keeps the database alive after close(): the file stays locked, and reopening it in the same process hangs. Details are in xqlsystems/duckdb-zarr#52.

    This doesn't change the proposal above. duckdb-zarr will publish the mapping as columns of read_zarr_groups (group, dims, table_name), and users create the views from them. The part that has to match between the two projects is still the node → schema → table mapping and the table names. xarray-sql can implement it as a real CatalogProvider.

  2. alxmrs commented on Oct 9, 2026

    @alxmrs
    MemberAuthor

    Two updates from implementing this in duckdb-zarr (xqlsystems/duckdb-zarr#52 for the tree, #56 for AnnData).

    The group parameter is group_path, not group. GROUP is a reserved word in DuckDB, so read_zarr(store, group := '...') doesn't parse. duckdb-zarr uses group_path :=, with the quoted alias "group" :=, the same pattern as its existing array_path / "array". In read_zarr_groups, the column is group_path and holds the DataTree path (/, /simulation/fine). A new schema_name column holds the schema each group maps to (main for the root, otherwise the path without the leading slash). Nothing else in the proposal above changed: one schema per group, tables named "_".join(dims), coordinates inherited from ancestors.

    AnnData: obs and var are not nodes after all. Real queries against the first version read badly: pbmc.obs.obs and pbmc.var.var, joined to X through an _index column. AnnData stores now map as follows (design decision 9 in duckdb-zarr):

    Element Dimensions Table
    X obs, var main.obs_var
    obs columns obs main.obs
    var columns var main.var
    layers/* obs, var layers.obs_var
    obsm/<k> obs, <k>_component obsm.obs_<k>_component
    obsp/* obs_i, obs_j obsp.obs_i_obs_j
    • obs and var are variables of the root node, next to X. This is the xarray Dataset view of an AnnData object: X(obs, var), cell_type(obs), gene_symbol(var).
    • Each data frame's index is the coordinate of its axis. The obs and var columns therefore hold cell and gene names in every table, inherited by layers, obsm and obsp.
    • Categoricals read as ENUMs, and sparse matrices read as one row per stored entry.

    If xarray-sql reads AnnData stores too, matching these rules would give the two projects the same tables. Examples of the queries this gives are in xqlsystems/duckdb-zarr#40.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions