Repository navigation
Shared design for DataTree tables with duckdb-zarr: one schema per node #262
Description
Activity
Update from the duckdb-zarr side: duckdb-zarr won't mount a store as a catalog itself, at least for now.
We spiked a
CALL zarr_attach(store, alias)that creates one schema per node and one view per dimension group. The DuckDB 1.5.5 C API has no storage-extension hook and can't run SQL on the calling connection, so the procedure needs a second connection that it opens at load and keeps. That connection keeps the database alive afterclose(): the file stays locked, and reopening it in the same process hangs. Details are in xqlsystems/duckdb-zarr#52.This doesn't change the proposal above. duckdb-zarr will publish the mapping as columns of
read_zarr_groups(group,dims,table_name), and users create the views from them. The part that has to match between the two projects is still the node → schema → table mapping and the table names. xarray-sql can implement it as a realCatalogProvider.Two updates from implementing this in duckdb-zarr (xqlsystems/duckdb-zarr#52 for the tree, #56 for AnnData).
The group parameter is
group_path, notgroup.GROUPis a reserved word in DuckDB, soread_zarr(store, group := '...')doesn't parse. duckdb-zarr usesgroup_path :=, with the quoted alias"group" :=, the same pattern as its existingarray_path/"array". Inread_zarr_groups, the column isgroup_pathand holds the DataTree path (/,/simulation/fine). A newschema_namecolumn holds the schema each group maps to (mainfor the root, otherwise the path without the leading slash). Nothing else in the proposal above changed: one schema per group, tables named"_".join(dims), coordinates inherited from ancestors.AnnData:
obsandvarare not nodes after all. Real queries against the first version read badly:pbmc.obs.obsandpbmc.var.var, joined toXthrough an_indexcolumn. AnnData stores now map as follows (design decision 9 in duckdb-zarr):Element Dimensions Table Xobs,varmain.obs_varobscolumnsobsmain.obsvarcolumnsvarmain.varlayers/*obs,varlayers.obs_varobsm/<k>obs,<k>_componentobsm.obs_<k>_componentobsp/*obs_i,obs_jobsp.obs_i_obs_jobsandvarare variables of the root node, next toX. This is the xarrayDatasetview of an AnnData object:X(obs, var),cell_type(obs),gene_symbol(var).- Each data frame's index is the coordinate of its axis. The
obsandvarcolumns therefore hold cell and gene names in every table, inherited bylayers,obsmandobsp. - Categoricals read as
ENUMs, and sparse matrices read as one row per stored entry.
If xarray-sql reads AnnData stores too, matching these rules would give the two projects the same tables. Examples of the queries this gives are in xqlsystems/duckdb-zarr#40.
duckdb-zarr is adding support for Zarr stores with nested groups, such as those written by
xarray.DataTree.to_zarr(). We'd like both projects to turn the same store into the same tables, with the same names, so this issue proposes a shared design. It builds on #82.The duckdb-zarr version is decision 8 in xqlsystems/duckdb-zarr#52 (design only, no code yet).
Proposal
DataTreemodel:simulation/coarseandsimulation/finecan both havefoo(x, y)with different lengths ofx.default_table_name("_".join(dims)in dimension order,scalarfor none) andresolve_table_names(overrides keyed by the dims tuple, case-insensitive collision check), applied per node.mainin DuckDB,publicin DataFusion.sim."simulation/fine".x_y.CatalogProviderwhoseSchemaProviders are the nodes, as suggested in Support for xarray-datatrees? #82.dt["simulation/fine"].to_dataset()gives.read_zarr(store, group := 'simulation/fine'), matchingxr.open_zarr(store, group=...). With nogroup, only the root node is read.Questions for xarray-sql
from_datatreeAPI? Or would you rather have flat table names in one schema (for examplesimulation_fine__x_y)?table_namesoverrides for a tree be keyed by(node_path, dims), with the per-nodedimskeys kept as they are?Datasetthat would make inherited coordinates hard?Why now
The immediate user is AnnData (xqlsystems/duckdb-zarr#40). An AnnData store is a tree:
obsandvarare nodes, andX,layersandobsmshare their axes. The AnnData-specific parts come after this: naming its axes and decoding its categorical, nullable and sparse groups. Agreeing on the tree mapping first means both projects expose AnnData the same way.