Skip to content

Latest commit

 

History

173 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ctest format

tinylamb

A simple implementation of RDBMS.

SQL frontend

The tinylamb executable reads a SQL script from standard input and executes every statement as one implicit transaction. The pinned GoogleSQL parser produces its parser AST, a visitor converts that tree directly to tinylamb expressions and relational query nodes, and the resulting plan is executed by tinylamb.

cmake -S . -B build
cmake --build build -j
printf 'CREATE TABLE t (k INT64, v VARCHAR); INSERT INTO t VALUES (1, %s);\nSELECT * FROM t WHERE k = 1;\n' "'hello'" | ./build/tinylamb /tmp/demo.db

tinylamb creates the database file on first use; it bootstraps only the catalog, so create your tables with DDL before querying.

CMake downloads the pinned GoogleSQL execute_query release and verifies its SHA-256 checksum. Bazel is not required. GoogleSQL AST mode is required by the SQL executable; configuring with -DTINYLAMB_ENABLE_GOOGLESQL=OFF still builds every library, executable and test binary, but the SQL frontend becomes unavailable at runtime (statements fail to parse and frontend-driven tests skip themselves).

The SQL execution path supports the TPC-C transaction query shapes (excluding stored procedures) and all 22 TPC-H queries. This includes inner and left joins, derived tables, CTEs, correlated scalar/IN/EXISTS subqueries, GROUP BY/HAVING, nested and distinct aggregates, CASE, LIKE, BETWEEN, date intervals and extraction, ORDER BY, LIMIT/OFFSET, INSERT, UPDATE, and DELETE. End-to-end coverage lives in query/sql_engine_tpcc_test.cpp and query/sql_engine_tpch_test.cpp.

EXPLAIN returns the selected physical execution strategy without running the query. EXPLAIN ANALYZE executes it and adds the actual row count plus total planning and execution time to the plan dump. Both forms work through standard input and PostgreSQL simple-query clients such as psql.

EXPLAIN SELECT * FROM lineitem WHERE l_orderkey = 1;
EXPLAIN ANALYZE SELECT l_returnflag, COUNT(*) FROM lineitem
  GROUP BY l_returnflag;

The TPC-H benchmark takes a scale factor, builds the pinned TPC-H DBGEN tool, generates all eight official-distribution .tbl files, validates their cardinalities, loads them with their SQL types, and executes all 22 queries. It prints each runtime profile followed by a slowest-first summary. DBGEN is only downloaded when this explicit benchmark target is built.

cmake --build build -j --target tinylamb_tpch_benchmark
./build/tinylamb_tpch_benchmark /var/tmp/tinylamb-tpch-sf1/database \
  --scale-factor 1 --data-dir /var/tmp/tinylamb-tpch-sf1/data

--reuse-database skips schema creation and loading for a previously loaded database, and --query N runs one query while investigating a plan. --force only replaces the three tinylamb database files at the exact path supplied. Scale factors outside the official TPC-H set are accepted for quick engineering smoke tests but are reported as non-official.

PostgreSQL-compatible TCP server

tinylamb_server exposes the same SQL engine through PostgreSQL's version 3 wire protocol. It uses a non-blocking epoll event loop and accepts multiple TCP clients. The standard-input tinylamb executable remains available for scripts and debugging.

cmake --build build -j --target tinylamb_server
./build/tinylamb_server warehouse.db --host 127.0.0.1 --port 54321 \
  --read-workers 8

# In another terminal:
psql -X "host=127.0.0.1 port=54321 user=tinylamb dbname=warehouse sslmode=prefer"

The server supports PostgreSQL startup negotiation, trust authentication, simple queries, text result rows and NULLs, command tags, error responses, and BEGIN/COMMIT/ROLLBACK. Autocommit SELECT messages run on a worker pool (statements that start with WITH or EXPLAIN take the normal execution path because they can wrap writes), so independent clients execute read-only transactions in parallel without blocking the epoll I/O loop. The default worker count is the detected hardware concurrency and can be set with --read-workers. TLS/GSS encryption requests are declined and clients continue over plain TCP. The default bind address is therefore 127.0.0.1; binding to a non-loopback address should only be done on a trusted network. PostgreSQL extended-query messages, COPY, catalog compatibility for psql backslash commands such as \\dt, and PostgreSQL user/password management are not implemented yet.

Read scaling can be measured with the bundled PostgreSQL-protocol benchmark client after creating a table named read_scale with an integer id column:

cmake --build build -j --target tinylamb_pg_read_benchmark
./build/tinylamb_pg_read_benchmark --host 127.0.0.1 --port 54321 \
  --clients 8 --warmup 3 --seconds 15 --database warehouse

Cascades optimizer

Logical joins are explored in a Cascades-style memo rather than by the old subset dynamic-programming loop. Scalar normalization, logical transformations, and physical implementations are independent named rule sets, each with a small C++ pattern DSL and add/remove APIs. See docs/cascades_optimizer.md for the architecture and extension examples.

Call path: SqlEngine builds QueryData, runs Rewrite, then calls Optimizer::Optimize (plan/optimizer.cpp). That entry point normalizes predicates, constructs a Cascades Search, and delegates join/order enumeration to plan/cascades.cpp via RuleSet / ImplementationRuleSet. There is no second standalone DP optimizer; optimizer.cpp is the façade and post-processing layer (aggregates, residual predicates).

Statement AST: SQL DML/SELECT shapes live in query/statement.hpp (executor-side code includes it directly).

Layer libraries: tinylamb_core is an INTERFACE aggregate over tinylamb_common → … → tinylamb_executor (see CMakeLists.txt). The layer map, ownership table, and include rules are documented in ARCHITECTURE.md; dependency direction is enforced by python3 scripts/check_layering.py.

Docs index: WAL/page format, lock order, recovery, checksum, durability, and Value storage notes are under docs/. See CONTRIBUTING.md for error-handling policy.

TPC-C workload benchmark

The benchmark client takes a TPC-C scale factor W (warehouse count, default

  1. and loads the Clause 4.3 population: 10 districts per warehouse, 3,000 customers per district, 100,000 items, and 3,000 initial orders per district (the newest 900 stay in NEW-ORDER). It then runs 10 terminals per warehouse with the standard 45/43/4/4/4 mix, NURand customer/item/last-name selection, 1% New-Order unused-item rollback, and 15% remote Payment.
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j --target tinylamb_tpcc_benchmark
./build/tinylamb_tpcc_benchmark /tmp/tinylamb-tpcc.db --scale-factor 1
./build/tinylamb_tpcc_benchmark /tmp/tinylamb-tpcc.db --sf 1 \
  --clients 10 --warmup 2 --seconds 10 --seed 20260819

Each SQL statement goes through the GoogleSQL frontend, tinylamb optimizer, executor, and transaction commit path. The client reports committed transactions per second (tps), SQL statements per second (sql_qps), and committed New-Order transactions per minute (new_order_tpm). Preflight checks all five transaction types, and the run fails unless every measured transaction executes at least one statement through SqlEngine (sql_path_gate=PASS). Think/keying time is omitted, so the number is not an audited TPC-C tpmC score. Scale factor 1 is a large load (on the order of 100k items and ~300k order lines). Unit tests use a reduced TpccScale::ForTest() population, not SF=1.

License

See ./LICENSE.txt

Copyright (c) 2023 KUMAZAKI Hiroki rintyo@gmail.com

About

Sample implementation of RDBMS.

Resources

Contributing

Stars

21 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages