Repository navigation
bug(server): forward service serializes connections on the SQLite store (two commits per TCP connection, rollback-journal mode) #3494
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 20, 2026 - added a commit that references this issue
on Sep 20, 2026 Follow-up with a before/after measurement using the proposed change, on a hosted GitHub Actions runner so it is reproducible without any particular deployment.
Setup:
ubuntu-latest(4 vCPU, ext4), gateway run directly on the runner with the Docker driver and a file-backed SQLite store, a sandbox running a loopback Python HTTP server (listen backlog raised to 1024 so the server itself is not the bottleneck),openshell forward servicein front of it, and a client that opens N simultaneous connections, each doingGET /withConnection: closeand reading 64 bytes. Both builds are the same source commit; the only difference is the store change from this issue (WAL +synchronous=NORMAL). Both sets ran back to back on the same runner with fresh state directories. The reproducer is on https://github.com/n1hility/OpenShell/tree/forward-sweep-repro: workflow.github/workflows/forward-sweep.yml(dispatch it with the defaultpatched_source=build; it pulls the public baseline images, builds the patched gateway from the checkout withdeploy/docker/Dockerfile.gha.gatewayin about 26 minutes, runs both sets back to back on one runner and writes the table to the job summary; no secrets or private images needed), scriptsscripts/forward-sweep/{run-set.sh,sweep.py,report.py}, and a "Running it by hand" section indeploy/docker/README-forge.mdfor any x86-64 Linux host with Docker: tworun-set.shinvocations withSET_NAME,GATEWAY_IMAGE,SUPERVISOR_IMAGE,CLI_IMAGE,OUT_DIR, thenreport.py. Any two gateway images can be compared; only the gateway differs between the sets. A run in build mode: https://github.com/rh-forge/openshell/actions/runs/35522232042 (2 to 9x on a quiet runner disk, same shape as below).Observed store mode: baseline
PRAGMA journal_mode = delete(no sidecars); patchedwal(-wal/-shmpresent). fdatasync of 4 KiB on this runner: p50 about 0.36 ms in both runs.Run with a quiet disk (max fsync 0.7 ms):
simultaneous connections baseline completed baseline wall (s) baseline mean per conn (ms) patched completed patched wall (s) patched mean per conn (ms) 1 1/1 0.005 5.3 1/1 0.004 4.3 6 6/6 0.037 14.4 6/6 0.009 6.9 16 16/16 0.108 37.3 16/16 0.020 13.8 32 32/32 0.122 72.0 24/32 0.048 22.1 64 51/64 0.335 126.6 55/64 0.081 42.3 10 sequential 10/10 0.046 4.6 10/10 0.027 2.7 Same setup on a runner whose disk showed fsync stalls up to 157 ms (the more realistic case for a VM on network storage):
simultaneous connections baseline completed baseline wall (s) baseline mean per conn (ms) patched completed patched wall (s) patched mean per conn (ms) 1 1/1 0.080 79.7 1/1 0.003 3.1 6 6/6 0.695 386 6/6 0.006 4.8 16 16/16 0.267 169 16/16 0.013 9.7 32 32/32 1.234 486 27/32 0.043 16.2 64 52/64 0.491 201 50/64 0.055 24.8 10 sequential 10/10 0.100 10.0 10/10 0.062 6.2 In that run the baseline gateway also logged sqlx "slow statement" warnings on the
ssh_sessionINSERT; the patched gateway logged none.Reading: the baseline's wall time grows with the number of simultaneous connections because the two commits per forwarded connection serialize on the rollback journal, and the size of the gap tracks the disk's commit latency (2.5 to 6x on the quiet disk, 10 to 100x when the disk stalls). With WAL the mean stays in the 3 to 40 ms range up to 64 simultaneous connections. The connections refused at 32 and 64 are the fixed 20-concurrent-per-sandbox cap in both builds; with the patched store they are refused because all connections arrive within a few milliseconds while slots are held for about 10 ms each, with the baseline because slots are held for hundreds of milliseconds. That cap is orthogonal to this change and worth its own discussion.
And the same measurement on the VM deployment from the original report, after upgrading its gateway to a build with the proposed change (2 vCPU VM, file-backed SQLite on a network block volume, sandboxed HTTPS service behind a forward, client in a pod on the VM's network). The store now reports
journal_mode = wal.simultaneous connections before: wall (s) / completed after: wall (s) / completed after: mean per conn (ms) 1 0.05 / 1 0.01 / 1 10 6 0.49 / 6 0.02 / 6 18 16 1.52 / 16 0.06 / 16 39-45 32 3.29 / 32 0.08 / 19 57-63 64 6.05 / 29 0.12 / 21 87 10 sequential mean 82 ms mean 7-19 ms Wall clock drops 25-50x and stops growing with N. At 32 and 64 the completions are now bounded by the 20-concurrent-per-sandbox cap, which refuses the excess within milliseconds; before the change the same connections were spread out by the store and mostly got through after seconds. So the cap becomes the visible limit once the store is fixed, which argues for treating it in a follow-up (configurable, or queue rather than refuse, or a token per forward process rather than per connection).
- addedstate:acceptedA maintainer decided OpenShell should pursue this issueA maintainer decided OpenShell should pursue this issueand removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 22, 2026 - added 2 commits that reference this issue
on Sep 26, 2026 - added 2 commits that reference this issue
on Oct 3, 2026
User Story
As an operator running a single-node OpenShell gateway on a small VM, with
openshell forward servicein front of HTTP services that run inside sandboxes, I want concurrent client connections through a forward to be served concurrently and cheaply, so that a client fetching a handful of records in parallel gets them in tens of milliseconds instead of seconds and does not see connections dropped.Problem Statement
Every TCP connection accepted by
openshell forward servicecosts about 50-100 ms before the first byte reaches the target, the wall-clock time for N simultaneous connections is linear in N, and beyond roughly 20 simultaneous connections the excess are closed by the forward instead of being queued.Reading the code, the forward path does two gateway store writes per TCP connection and the gateway's on-disk SQLite store runs in SQLite's default rollback-journal mode:
crates/openshell-cli/src/run.rsservice_forward_tcp: each accepted connection callsCreateSshSessionbefore theForwardTcpstream andRevokeSshSessionafter it. That is one INSERT and one UPDATE of anssh_sessionobject per connection.crates/openshell-server/src/persistence/sqlite.rsSqliteStore::connect: no journal or synchronous pragma is set, and sqlx 0.8 does not set one either, so a freshly created database runs withjournal_mode=deleteandsynchronous=FULL. Every commit pays several fsyncs and a writer blocks readers, so all the reads the forward path also does (fetch_and_authorize_sandbox, session validation) queue behind the writes. The existing comment inpersistence/tests.rs("sqlx 0.8 doesn't default to WAL ... actual production path today") and issue Control-plane read failures during sandbox SSH-session cleanup; delete acknowledgment can precede durable cleanup #2999 (observedjournal_mode=delete) confirm this is the shipped state.crates/openshell-server/src/grpc/sandbox.rsacquire_ssh_connection_slots: hard limit of 20 concurrent forward connections per sandbox (and 3 per token), introduced by fix(security): add SSH session token expiry, connection limits, and lifecycle cleanup #182 for SSH sessions. Connections queued behind the store are released in batches, so a burst exceeds the cap; the excess getRESOURCE_EXHAUSTED("sandbox SSH connection limit reached") and the CLI closes the client socket, which the client sees as an EOF or reset during its TLS handshake.Impact / Why This Matters
Any client that opens several connections to a forwarded service at once (a browser page loading N records, a connection pool warming up) serializes on the gateway store, and a burst above the cap loses connections outright. A dropped connection is worse than a slow one: the client has to detect it and retry, usually after a timeout. Current workarounds are to limit client concurrency to well under 20 and to retry dropped connections; neither removes the per-connection floor, which is 50-100 ms on a VM against a network-backed disk, and both are things every consumer of a forward has to know about.
Measured on a 2 vCPU Fedora 44 VM (gateway/CLI 0.0.110, Podman driver, file-backed SQLite on btrfs over a virtio network block volume; the relevant code is unchanged on
main), from a pod on the VM's network (TCP connect to the VM ~0.2 ms), against a sandboxed HTTPS service behind a forward:Concurrency sweep, N simultaneous connections each doing TLS handshake +
GET /+Connection: close, read 64 bytes:A second run 10 s later gave the same shape with 9 of 32 and 43 of 64 lost. Wall time is ~95-100 ms per connection at every N. Ten sequential single connections averaged 82 ms.
Control on the same VM, ten sequential bare TLS handshakes:
On the VM: the gateway database reports
PRAGMA journal_mode = delete, no-wal/-shmsidecars exist, and the forward units' journals contain 84 occurrences ofservice forward connection failed ... code: 'Some resource has been exhausted', message: "sandbox SSH connection limit reached"over 30 days, arriving in bursts (three within 7 ms).A benchmark on a copy of that database (100 x {INSERT; UPDATE} as autocommit statements, then a writer thread looping the same pair while a second connection times 100
SELECT count(*); note this copy ran on tmpfs, so the absolute commit costs are lower bounds without disk fsync):Even with no disk in the path, delete mode makes readers wait behind the writer and cuts write throughput 4x; with real fsync latency the per-commit cost grows to tens of milliseconds, which matches the ~95 ms per connection (two commits) seen through the forward.
Acceptance Criteria
openshell forward serviceno longer scales linearly with the number of simultaneous connections on a file-backed SQLite store; 32 simultaneous connections to a loopback HTTP server in a sandbox complete in well under a second on a small VM.-wal/-shmsidecars, and the backup implication.Out of scope for this issue but worth separate discussion: minting one session token per forwarded TCP connection (rather than per forward process), and making the per-sandbox connection cap configurable or turning a refusal into a queue.
Reproduction Steps
openshell sandbox create --name forward-probe -- sh -lc 'exec python3 -m http.server 63152 --bind 127.0.0.1'openshell forward service forward-probe --target-host 127.0.0.1 --target-port 63152 --local 127.0.0.1:43152127.0.0.1:43152, each sendingGET / HTTP/1.1withConnection: closeand reading the first bytes, for N in 1, 6, 16, 32, 64, and record wall time and how many connections complete. Repeat with ten sequential connections.sqlite3 <gateway-db> 'pragma journal_mode'(expectdelete).pragma journal_mode=walon the database file once, start it again and repeat step 4; the per-connection time drops and readers stop queueing.Environment
main.Logs
(84 occurrences in 30 days across the forward units on one VM, in bursts of several per 10 ms.)
Proposed change
Set
journal_mode=WALandsynchronous=NORMALfor on-disk SQLite stores inSqliteStore::connect, switching the file to WAL on a single connection before the pool opens (the mode change needs an exclusive lock thatbusy_timeoutcannot wait for), with tests and docs. A branch with this change, passing the persistence tests, is at https://github.com/n1hility/OpenShell/tree/sqlite-wal/jg; I will open it as a PR once this issue is triaged and I am vouched.