Skip to content

bug(server): forward service serializes connections on the SQLite store (two commits per TCP connection, rollback-journal mode) #3494

Description

@n1hility

User Story

As an operator running a single-node OpenShell gateway on a small VM, with openshell forward service in front of HTTP services that run inside sandboxes, I want concurrent client connections through a forward to be served concurrently and cheaply, so that a client fetching a handful of records in parallel gets them in tens of milliseconds instead of seconds and does not see connections dropped.

Problem Statement

Every TCP connection accepted by openshell forward service costs about 50-100 ms before the first byte reaches the target, the wall-clock time for N simultaneous connections is linear in N, and beyond roughly 20 simultaneous connections the excess are closed by the forward instead of being queued.

Reading the code, the forward path does two gateway store writes per TCP connection and the gateway's on-disk SQLite store runs in SQLite's default rollback-journal mode:

  • crates/openshell-cli/src/run.rs service_forward_tcp: each accepted connection calls CreateSshSession before the ForwardTcp stream and RevokeSshSession after it. That is one INSERT and one UPDATE of an ssh_session object per connection.
  • crates/openshell-server/src/persistence/sqlite.rs SqliteStore::connect: no journal or synchronous pragma is set, and sqlx 0.8 does not set one either, so a freshly created database runs with journal_mode=delete and synchronous=FULL. Every commit pays several fsyncs and a writer blocks readers, so all the reads the forward path also does (fetch_and_authorize_sandbox, session validation) queue behind the writes. The existing comment in persistence/tests.rs ("sqlx 0.8 doesn't default to WAL ... actual production path today") and issue Control-plane read failures during sandbox SSH-session cleanup; delete acknowledgment can precede durable cleanup #2999 (observed journal_mode=delete) confirm this is the shipped state.
  • crates/openshell-server/src/grpc/sandbox.rs acquire_ssh_connection_slots: hard limit of 20 concurrent forward connections per sandbox (and 3 per token), introduced by fix(security): add SSH session token expiry, connection limits, and lifecycle cleanup #182 for SSH sessions. Connections queued behind the store are released in batches, so a burst exceeds the cap; the excess get RESOURCE_EXHAUSTED ("sandbox SSH connection limit reached") and the CLI closes the client socket, which the client sees as an EOF or reset during its TLS handshake.

Impact / Why This Matters

Any client that opens several connections to a forwarded service at once (a browser page loading N records, a connection pool warming up) serializes on the gateway store, and a burst above the cap loses connections outright. A dropped connection is worse than a slow one: the client has to detect it and retry, usually after a timeout. Current workarounds are to limit client concurrency to well under 20 and to retry dropped connections; neither removes the per-connection floor, which is 50-100 ms on a VM against a network-backed disk, and both are things every consumer of a forward has to know about.

Measured on a 2 vCPU Fedora 44 VM (gateway/CLI 0.0.110, Podman driver, file-backed SQLite on btrfs over a virtio network block volume; the relevant code is unchanged on main), from a pod on the VM's network (TCP connect to the VM ~0.2 ms), against a sandboxed HTTPS service behind a forward:

Concurrency sweep, N simultaneous connections each doing TLS handshake + GET / + Connection: close, read 64 bytes:

N wall (s) completed mean per connection (ms) failures
1 0.05 1 54 0
6 0.49 6 264 0
16 1.52 16 772 0
32 3.29 32 1445 0
64 6.05 29 3159 35 (EOF/RST after ClientHello)

A second run 10 s later gave the same shape with 9 of 32 and 43 of 64 lost. Wall time is ~95-100 ms per connection at every N. Ten sequential single connections averaged 82 ms.

Control on the same VM, ten sequential bare TLS handshakes:

Path TLS handshake mean (ms)
gateway port directly, no forward 1.5
same VM, through a forward 88 (min 47, max 283)

On the VM: the gateway database reports PRAGMA journal_mode = delete, no -wal/-shm sidecars exist, and the forward units' journals contain 84 occurrences of service forward connection failed ... code: 'Some resource has been exhausted', message: "sandbox SSH connection limit reached" over 30 days, arriving in bursts (three within 7 ms).

A benchmark on a copy of that database (100 x {INSERT; UPDATE} as autocommit statements, then a writer thread looping the same pair while a second connection times 100 SELECT count(*); note this copy ran on tmpfs, so the absolute commit costs are lower bounds without disk fsync):

config ms per commit reader p95 (ms) with concurrent writer reader max (ms) writer iterations in 3 s
delete / FULL (current default) 0.08 18.4 53.5 18,815
wal / FULL 0.04 0.20 3.0 71,957
wal / NORMAL 0.02 0.17 0.4 71,581

Even with no disk in the path, delete mode makes readers wait behind the writer and cuts write throughput 4x; with real fsync latency the per-commit cost grows to tens of milliseconds, which matches the ~95 ms per connection (two commits) seen through the forward.

Acceptance Criteria

  • Connection setup time through openshell forward service no longer scales linearly with the number of simultaneous connections on a file-backed SQLite store; 32 simultaneous connections to a loopback HTTP server in a sandbox complete in well under a second on a small VM.
  • Reads on the gateway store are not blocked by concurrent writes (on-disk SQLite runs in WAL mode).
  • The behaviour is covered by a file-backed store test (journal mode after connect, existing rollback-journal files switched on connect, concurrent readers under a burst of insert-then-update writes).
  • Documentation states the SQLite durability setting, the -wal/-shm sidecars, and the backup implication.

Out of scope for this issue but worth separate discussion: minting one session token per forwarded TCP connection (rather than per forward process), and making the per-sandbox connection cap configurable or turning a refusal into a queue.

Reproduction Steps

  1. Run a gateway with the default SQLite store on a VM or host whose disk has non-trivial fsync latency (a cloud block volume is enough).
  2. Create a sandbox running a loopback HTTP server, for example:
    openshell sandbox create --name forward-probe -- sh -lc 'exec python3 -m http.server 63152 --bind 127.0.0.1'
  3. openshell forward service forward-probe --target-host 127.0.0.1 --target-port 63152 --local 127.0.0.1:43152
  4. From the same host, open N simultaneous TCP connections to 127.0.0.1:43152, each sending GET / HTTP/1.1 with Connection: close and reading the first bytes, for N in 1, 6, 16, 32, 64, and record wall time and how many connections complete. Repeat with ten sequential connections.
  5. Compare with N connections to the HTTP server without the forward (from inside the sandbox) and with sqlite3 <gateway-db> 'pragma journal_mode' (expect delete).
  6. Optional confirmation: stop the gateway, run pragma journal_mode=wal on the database file once, start it again and repeat step 4; the per-connection time drops and readers stop queueing.

Environment

  • OpenShell: measured on 0.0.110 gateway and CLI; the store, forward and connection-cap code is the same on main.
  • OS: Fedora Linux 44 (Cloud), kernel 6.19, 2 vCPU / 4 GiB VM.
  • Runtime/deployment: Podman compute driver, single gateway, default file-backed SQLite store on btrfs over a virtio network block volume.

Logs

WARN service forward connection failed peer=<ip>:<port> error=code: 'Some resource has been exhausted', message: "sandbox SSH connection limit reached"

(84 occurrences in 30 days across the forward units on one VM, in bursts of several per 10 ms.)

Proposed change

Set journal_mode=WAL and synchronous=NORMAL for on-disk SQLite stores in SqliteStore::connect, switching the file to WAL on a single connection before the pool opens (the mode change needs an exclusive lock that busy_timeout cannot wait for), with tests and docs. A branch with this change, passing the persistence tests, is at https://github.com/n1hility/OpenShell/tree/sqlite-wal/jg; I will open it as a PR once this issue is triaged and I am vouched.

Activity

  1. n1hility commented on Sep 20, 2026

    @n1hility
    ContributorAuthor

    Follow-up with a before/after measurement using the proposed change, on a hosted GitHub Actions runner so it is reproducible without any particular deployment.

    Setup: ubuntu-latest (4 vCPU, ext4), gateway run directly on the runner with the Docker driver and a file-backed SQLite store, a sandbox running a loopback Python HTTP server (listen backlog raised to 1024 so the server itself is not the bottleneck), openshell forward service in front of it, and a client that opens N simultaneous connections, each doing GET / with Connection: close and reading 64 bytes. Both builds are the same source commit; the only difference is the store change from this issue (WAL + synchronous=NORMAL). Both sets ran back to back on the same runner with fresh state directories. The reproducer is on https://github.com/n1hility/OpenShell/tree/forward-sweep-repro: workflow .github/workflows/forward-sweep.yml (dispatch it with the default patched_source=build; it pulls the public baseline images, builds the patched gateway from the checkout with deploy/docker/Dockerfile.gha.gateway in about 26 minutes, runs both sets back to back on one runner and writes the table to the job summary; no secrets or private images needed), scripts scripts/forward-sweep/{run-set.sh,sweep.py,report.py}, and a "Running it by hand" section in deploy/docker/README-forge.md for any x86-64 Linux host with Docker: two run-set.sh invocations with SET_NAME, GATEWAY_IMAGE, SUPERVISOR_IMAGE, CLI_IMAGE, OUT_DIR, then report.py. Any two gateway images can be compared; only the gateway differs between the sets. A run in build mode: https://github.com/rh-forge/openshell/actions/runs/35522232042 (2 to 9x on a quiet runner disk, same shape as below).

    Observed store mode: baseline PRAGMA journal_mode = delete (no sidecars); patched wal (-wal/-shm present). fdatasync of 4 KiB on this runner: p50 about 0.36 ms in both runs.

    Run with a quiet disk (max fsync 0.7 ms):

    simultaneous connections baseline completed baseline wall (s) baseline mean per conn (ms) patched completed patched wall (s) patched mean per conn (ms)
    1 1/1 0.005 5.3 1/1 0.004 4.3
    6 6/6 0.037 14.4 6/6 0.009 6.9
    16 16/16 0.108 37.3 16/16 0.020 13.8
    32 32/32 0.122 72.0 24/32 0.048 22.1
    64 51/64 0.335 126.6 55/64 0.081 42.3
    10 sequential 10/10 0.046 4.6 10/10 0.027 2.7

    Same setup on a runner whose disk showed fsync stalls up to 157 ms (the more realistic case for a VM on network storage):

    simultaneous connections baseline completed baseline wall (s) baseline mean per conn (ms) patched completed patched wall (s) patched mean per conn (ms)
    1 1/1 0.080 79.7 1/1 0.003 3.1
    6 6/6 0.695 386 6/6 0.006 4.8
    16 16/16 0.267 169 16/16 0.013 9.7
    32 32/32 1.234 486 27/32 0.043 16.2
    64 52/64 0.491 201 50/64 0.055 24.8
    10 sequential 10/10 0.100 10.0 10/10 0.062 6.2

    In that run the baseline gateway also logged sqlx "slow statement" warnings on the ssh_session INSERT; the patched gateway logged none.

    Reading: the baseline's wall time grows with the number of simultaneous connections because the two commits per forwarded connection serialize on the rollback journal, and the size of the gap tracks the disk's commit latency (2.5 to 6x on the quiet disk, 10 to 100x when the disk stalls). With WAL the mean stays in the 3 to 40 ms range up to 64 simultaneous connections. The connections refused at 32 and 64 are the fixed 20-concurrent-per-sandbox cap in both builds; with the patched store they are refused because all connections arrive within a few milliseconds while slots are held for about 10 ms each, with the baseline because slots are held for hundreds of milliseconds. That cap is orthogonal to this change and worth its own discussion.

  2. n1hility commented on Sep 20, 2026

    @n1hility
    ContributorAuthor

    And the same measurement on the VM deployment from the original report, after upgrading its gateway to a build with the proposed change (2 vCPU VM, file-backed SQLite on a network block volume, sandboxed HTTPS service behind a forward, client in a pod on the VM's network). The store now reports journal_mode = wal.

    simultaneous connections before: wall (s) / completed after: wall (s) / completed after: mean per conn (ms)
    1 0.05 / 1 0.01 / 1 10
    6 0.49 / 6 0.02 / 6 18
    16 1.52 / 16 0.06 / 16 39-45
    32 3.29 / 32 0.08 / 19 57-63
    64 6.05 / 29 0.12 / 21 87
    10 sequential mean 82 ms mean 7-19 ms

    Wall clock drops 25-50x and stops growing with N. At 32 and 64 the completions are now bounded by the 20-concurrent-per-sandbox cap, which refuses the excess within milliseconds; before the change the same connections were spread out by the store and mostly got through after seconds. So the cap becomes the visible limit once the store is fixed, which argues for treating it in a follow-up (configurable, or queue rather than refuse, or a token per forward process rather than per connection).

  3. added
    state:acceptedA maintainer decided OpenShell should pursue this issue
    and removed
    state:triage-neededOpened without agent diagnostics and needs triage
    on Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:acceptedA maintainer decided OpenShell should pursue this issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions