Skip to content

feat: add lazy Arrow and database viewing - #1804

Open
Fred-Wu wants to merge 3 commits into
REditorSupport:mainfrom
Fred-Wu:feature/lazy-data-view-final
Open

Fred-Wu wants to merge 3 commits into
REditorSupport:mainfrom
Fred-Wu:feature/lazy-data-view-final

Conversation

@Fred-Wu

@Fred-Wu Fred-Wu commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Closes #1785

This might be a big change to support both Arrow's FileSystemDataset and SQL Server in the Data Viewer. I submit it here for review and discussion. It is not necessarily intended to adopt both of them. The implementations are more complicated than I initially thought, especially when performance is one part of the goals. Nevertheless, the following decisions are mine, with GPT used to generate and iterate on the implementation here.

Summary

Add lazy/on-demand Data Viewer support for large Arrow and DBI-backed data sources without loading the full dataset into memory.

Arrow and DBI now share the same general paging model:

  • fetch data on demand while scrolling;
  • use 1,000-row blocks for normal access and 5,000-row blocks for sorted results;
  • retain up to 20,000 fetched rows in the cache;
  • reuse cached rows for backward scrolling;
  • invalidate the relevant cache when filtering, sorting, or column projection changes;
  • fetch only columns required by the current page;
  • apply filtering and sorting to the complete data source rather than only loaded rows.
  • Preserve source ordering where applicable, while allowing Data Viewer sorting to take precedence over the same columns (e.g. dplyr::arrange()).
  • clean up backend resources when the viewer closes.

The existing in-memory Data Viewer behaviour remains unchanged.

Changes

Arrow

  • Reuse a persistent Arrow reader for sequential forward scrolling.
  • Keep filtered/sorted source-row indices so arbitrary display positions can be mapped back to source rows.
  • Maintain separate source-row and filtered/sorted display caches.
  • Use fragment-aware access for multi-file datasets.
  • Use Parquet row-group access for distant or sparse requests.
  • Preserve partition columns when reading individual files or row groups.
  • Preserve Arrow-specific column types during paging, including integer64, dates, datetimes, durations, and nested columns.

DBI

  • Add lazy viewing for tbl_sql objects, currently targeting SQL Server.
  • Push filtering and sorting to SQL Server.
  • Reuse an open DBI result with dbFetch() for continuous forward scrolling.
  • Use OFFSET to reposition the result for jumps or backward requests outside the cache.
  • Avoid placing source ORDER BY clauses inside SQL Server subqueries where they are invalid.

Other changes

  • sess/R/handlers.R

    • Route lazy Arrow and DBI objects through the Data Viewer.
    • Add shared paging, filtering, sorting, column-projection, and row-index handling used by the lazy viewers.
    • Dispose backend resources when a Data Viewer is closed.
  • sess/R/server.R

    • Add Data Viewer init, page, and dispose RPC handling.
    • Allow long-running page requests to be interrupted.
    • Report active Data Viewer reads so the extension can interrupt the correct request.
    • Ensure sess continues processing subsequent requests after an interrupted Data Viewer fetch.
  • src/session.ts

    • Add the dynamic Data Viewer request bridge.
    • Send selected column fields with page requests.
    • Serialise page requests and discard queued requests after the viewer closes.
    • Interrupt an active Data Viewer read when its panel is closed.
    • Remove the normal request timeout for page and dispose operations while retaining the initialisation timeout.
  • sess/DESCRIPTION

    • Add the optional dependencies required by lazy Arrow and DBI Data Viewer support.
  • Data Viewer tests

    • Extend the existing Data Viewer tests and add coverage for lazy paging, column projection, caching, filtering, sorting, interruption, disposal, Arrow fragment/Parquet access, and DBI behaviour.

Arrow and DBI use backend-specific strategies under the same Data Viewer behaviour: Arrow keeps source-row mappings for efficient random access, while DBI lets SQL Server perform filtering and sorting and streams the resulting rows through an open database result.

Validation

Added regression tests covering:

  • Arrow sequential and random paging;
  • multi-file fragment access and Parquet row groups;
  • filtering, sorting, caching, and column projection;
  • DBI SQL generation and source ordering;
  • DBI cursor reuse and repositioning;
  • viewer interruption and disposal;
  • date, datetime, logical, formatted, and integer64 columns.

@Fred-Wu
Fred-Wu marked this pull request as ready for review September 30, 2026 14:02
@Fred-Wu
Fred-Wu requested review from eitsupi, grantmcdermott, randy3k and renkun-ken and removed request for eitsupi and renkun-ken September 30, 2026 14:03

@eitsupi eitsupi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have a few high-level questions/concerns about the current approach.

SQL support

Is there a reason why the SQL abstraction cannot be left to dbplyr?
Maintaining SQL Server-specific query generation in vscode-R looks potentially fragile to me.

arrow dplyr query

I am also a little concerned about the amount of work being done when filtering or sorting.
For a complex lazy dplyr query, it looks like some expensive parts of the query could be evaluated repeatedly as the viewer state changes.
I may be particularly sensitive to this because I have contributed to the Arrow R package a number of times.

Data conversion

Arrow has a richer type system than R data frames, so it feels unfortunate to convert Arrow data back into R data frames as part of the viewing pipeline.
Would it make sense to preserve Arrow data through to the JS side instead?

Embedding DuckDB would be one relatively easy way to process Arrow data there, although of course DuckDB does not support every Arrow type either.

@Fred-Wu

Fred-Wu commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

I have a few high-level questions/concerns about the current approach.

Thanks for the review @eitsupi. Those concerns are valid. The main design choices are around the data viewer layer for forward/backward scrolling, including those operations after filtering and sorting and keeping them responsive and efficient.

For SQL Server

You are right; some SQL Server-specific abstractions could essentially be handed down to dbplyr, such as sorting, column projection, or metadata collection. What should be retained is the data viewer execution layer. For example, forward row fetching/paging uses dbFetch() during scrolling while caching paged rows for backward scrolling. For random jumps outside the cached pages, the viewer would need to decide when to execute again at the new position. But I think random-position execution could use dbplyr as well.

One special case is tbl_sql that is already arrange()ed. When that lazy query is rendered as a subquery, its ORDER BYstatement cannot be relied on to preserve the result order because dbplyr drops that column sorting from the subquery. To preserve the original lazy query semantic, the viewer needs to capture the columns to be sorted and reapply them to the outer query. If one, later, uses data viewer to sort columns, those orderings should take precedence.

arrow_dplyr_query

The major work on the arrow side is actually around FileSystemDataset and how the data viewer layer could execute on it. Support for arrow_dplyr_query is naturally built upon it. This is similar to the SQL side that the main question is when a lazy query should actually be executed and how much of the previous execution can be reused.

The more expensive re-execution can happen when the viewer filter or sort changes, and the requested page is then fetched separately. So for a complex arrow_dplyr_query, some operations may be evaluated more than once.

I thought about disabling filtering and sorting for Arrow and SQL, but the initial testing on large dummy datasets (both single-file and multi-file) and on a SQL Server was not bad without using lazy queries.

Any improvement ideas are welcome, especially around whether the evaluated state can be reused without losing the random access and backward scrolling. Maybe we can also have a list of operations that can be supported and disabled.

As for Arrow data types, the current Data Viewer is designed around data types supported by R. Preserving other Arrow data types may require broader changes, which are outside the original scope of this work.

@Fred-Wu

Fred-Wu commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

I think I may just disable filtering and sorting for lazy queries such as arrow_dplyr_query as if they are the results that want to be viewed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

(feat) Additional Arrow support

2 participants