Description
Implement the initial Expressif Tree-sitter grammar covering the fundamental syntax required to parse independently testable expressions:
- function calls
- positional arguments
- function pipelines
- numeric literals
- boolean literals
- quoted literals
- temporal literals
- open expressions
- closed expressions
- root expressions
This issue establishes the basic syntax tree on which the remaining Expressif grammar will be built.
Design principles
The grammar must remain independent from the Expressif function catalogue. The parser must not know whether a callable exists, whether it is a function/predicate/accumulator/transformation, what argument types it accepts, or how many arguments it accepts.
The syntax should avoid constructs whose interpretation depends on semantic context. Consequently:
- bare alphabetic names are function calls;
- unquoted textual literals are not supported;
- textual values must be quoted;
- temporal values use an explicit
#"..." form;
- booleans use
#true and #false;
- numeric values remain numeric literals.
Syntax model
RootExpressionSyntax
├── OpenExpressionSyntax
└── ClosedExpressionSyntax
ExpressionSyntax
└── FunctionCallSyntax
ArgumentSyntax
└── PositionalArgumentSyntax
ValueSyntax
├── NumericLiteralSyntax
├── BooleanLiteralSyntax
├── QuotedLiteralSyntax
└── TemporalLiteralSyntax
├── DateLiteralSyntax
├── DateTimeLiteralSyntax
└── TimeLiteralSyntax
Exact Tree-sitter rule names may follow Tree-sitter naming conventions, but the resulting CST should preserve these distinctions.
Function names
A function name consists of one or more alphabetic segments separated by a single hyphen.
Valid examples:
lower
LOWER
text-to-lower
foo-BAr-foo
true
false
Invalid examples:
1foo
foo1
fo1o
-foo
foo-
foo--bar
@foo
Function names are case-preserving at syntax level. Name normalization, aliases and callable resolution belong to the semantic binding layer.
Function calls
Zero-argument functions may omit parentheses:
or use empty parentheses:
Functions may contain positional arguments:
add(5)
round(2)
example(10, 20)
remove-chars("a")
The parser must not validate existence or arity. Therefore unknown, unknown(), unknown(5) and unknown(5, 10) are all syntactically valid.
Numeric literals
Support numeric literals such as:
They produce NumericLiteralSyntax. Preserve the source representation; conversion to runtime numeric types belongs to the binding layer.
Boolean literals
Boolean literals use explicit syntax:
and produce BooleanLiteralSyntax.
Bare true and false are function calls, not boolean literals.
Initially support exactly lowercase #true and #false unless case-insensitivity is deliberately added to the language definition.
Quoted literals
Textual values must be quoted.
Valid:
"foo"
""
"foo bar"
"foo , bar"
`foo`
` foo bar `
`foo , bar`
`(foo)`
Escaped double quotes must be supported:
The CST must retain the original quoting style.
Unquoted textual parameters such as remove-chars(a), append(*) or culture(fr-fr) are intentionally no longer textual literals. Their quoted equivalents must be used instead.
Temporal literals
Temporal values use the explicit #"..." syntax and are lexically distinct from ordinary quoted text.
Support three forms:
#"2025-12-17"
#"2025-12-17T14:30:00"
#"2025-12-17 14:30:00"
#"14:30:00"
They map respectively to:
DateLiteralSyntax
DateTimeLiteralSyntax
DateTimeLiteralSyntax
TimeLiteralSyntax
Accepted lexical formats are strictly:
yyyy-MM-dd
yyyy-MM-ddTHH:mm:ss
yyyy-MM-dd HH:mm:ss
HH:mm:ss
T and a space are alternative separators for the same DateTimeLiteralSyntax. The CST should preserve which separator was used.
The parser should validate the structural format, but semantic calendar validity such as February 30 may remain a later validation concern.
Ordinary quoted text that happens to look temporal remains text:
"2025-12-17" // QuotedLiteralSyntax
#"2025-12-17" // DateLiteralSyntax
Positional arguments
Arguments may initially contain any value implemented by this issue:
add(5)
between(5, 10)
foo(#true)
foo(#false)
foo("bar")
foo(`bar`)
foo(#"2025-12-17")
Later issues will extend arguments with variables/references, compound values, expressions, named arguments and spread arguments. Do not introduce function-name-specific argument grammar.
Open expressions
An open expression consists of one or more function calls chained with |:
lower
lower()
lower | trim
add(5) | multiply(2)
remove-chars("a") | upper
Whitespace around | is insignificant.
Closed expressions
A closed expression starts with a value and may optionally continue with a function pipeline:
10
10 | add(5)
#true
#true | some-function
"foo"
"foo" | lower
#"2025-12-17"
#"2025-12-17" | some-function
Additional source types such as variables, arrays, tuples, records, references and intervals will be introduced by later issues.
Root expressions
The parser entry point must classify the input as either an open or closed expression.
Examples:
lower → OpenExpressionSyntax
lower() | trim → OpenExpressionSyntax
true → OpenExpressionSyntax
10 → ClosedExpressionSyntax
10 | add(5) → ClosedExpressionSyntax
#true → ClosedExpressionSyntax
"foo" → ClosedExpressionSyntax
#"2025-12-17" → ClosedExpressionSyntax
There should be no ambiguity between a bare function name and a literal.
Whitespace
Accept insignificant surrounding whitespace, including:
lower
lower
lower | trim
add( 5 )
add(5, 10)
add(5 , 10)
Whitespace inside quoted literals remains significant. The space separator inside a datetime literal is part of the literal format.
Breaking changes from current Expressif syntax
The existing Expressif parser accepts unquoted textual values such as:
remove-chars(a)
append(*)
foo(fr-fr)
This grammar deliberately removes that ambiguity. Equivalent syntax becomes:
remove-chars("a")
append("*")
foo("fr-fr")
Boolean literals are explicitly written as #true and #false. Temporal literals are explicitly written using #"...".
These changes ensure that bare alphabetic tokens have one syntactic meaning: function calls.
Out of scope
Do not implement or special-case:
- variables (
@foo)
- property references (
[foo])
- index references (
#1)
- arrays
- tuples
- records
- intervals
- parameterized expressions
- named arguments (
name := ...)
- spread arguments (
...)
- field-access shorthand (
.foo)
- tuple-projection shorthand (
$0)
- map shorthand (
|>)
- unary negation
- binary operators
- grouped logical expressions
- semantic callable resolution
- callable arity validation
- callable parameter-type validation
Tests
Add Tree-sitter corpus tests covering at minimum:
- zero-argument function without parentheses;
- zero-argument function with parentheses;
- hyphen-separated and mixed-case function names;
- invalid function names;
- one and multiple positional arguments;
- positive, negative and decimal numeric literals;
#true and #false;
- bare
true and false parsed as function calls;
- double-quoted and backtick-quoted text;
- empty quoted text;
- quoted text containing spaces, commas and parentheses;
- escaped double quotes;
- valid date literals;
- valid datetime literals using
T and space;
- valid time literals;
- ordinary quoted temporal-looking text remaining a text literal;
- malformed temporal formats;
- one-function and multi-function open expressions;
- numeric, boolean, text and temporal closed expressions;
- closed expressions followed by pipelines;
- whitespace variations;
- malformed function calls;
- rejection of unquoted textual arguments.
Tests should assert syntax-tree structure, not only parse success.
Description
Implement the initial Expressif Tree-sitter grammar covering the fundamental syntax required to parse independently testable expressions:
This issue establishes the basic syntax tree on which the remaining Expressif grammar will be built.
Design principles
The grammar must remain independent from the Expressif function catalogue. The parser must not know whether a callable exists, whether it is a function/predicate/accumulator/transformation, what argument types it accepts, or how many arguments it accepts.
The syntax should avoid constructs whose interpretation depends on semantic context. Consequently:
#"..."form;#trueand#false;Syntax model
Exact Tree-sitter rule names may follow Tree-sitter naming conventions, but the resulting CST should preserve these distinctions.
Function names
A function name consists of one or more alphabetic segments separated by a single hyphen.
Valid examples:
Invalid examples:
Function names are case-preserving at syntax level. Name normalization, aliases and callable resolution belong to the semantic binding layer.
Function calls
Zero-argument functions may omit parentheses:
or use empty parentheses:
Functions may contain positional arguments:
The parser must not validate existence or arity. Therefore
unknown,unknown(),unknown(5)andunknown(5, 10)are all syntactically valid.Numeric literals
Support numeric literals such as:
They produce
NumericLiteralSyntax. Preserve the source representation; conversion to runtime numeric types belongs to the binding layer.Boolean literals
Boolean literals use explicit syntax:
and produce
BooleanLiteralSyntax.Bare
trueandfalseare function calls, not boolean literals.Initially support exactly lowercase
#trueand#falseunless case-insensitivity is deliberately added to the language definition.Quoted literals
Textual values must be quoted.
Valid:
Escaped double quotes must be supported:
The CST must retain the original quoting style.
Unquoted textual parameters such as
remove-chars(a),append(*)orculture(fr-fr)are intentionally no longer textual literals. Their quoted equivalents must be used instead.Temporal literals
Temporal values use the explicit
#"..."syntax and are lexically distinct from ordinary quoted text.Support three forms:
They map respectively to:
Accepted lexical formats are strictly:
Tand a space are alternative separators for the sameDateTimeLiteralSyntax. The CST should preserve which separator was used.The parser should validate the structural format, but semantic calendar validity such as February 30 may remain a later validation concern.
Ordinary quoted text that happens to look temporal remains text:
Positional arguments
Arguments may initially contain any value implemented by this issue:
Later issues will extend arguments with variables/references, compound values, expressions, named arguments and spread arguments. Do not introduce function-name-specific argument grammar.
Open expressions
An open expression consists of one or more function calls chained with
|:Whitespace around
|is insignificant.Closed expressions
A closed expression starts with a value and may optionally continue with a function pipeline:
Additional source types such as variables, arrays, tuples, records, references and intervals will be introduced by later issues.
Root expressions
The parser entry point must classify the input as either an open or closed expression.
Examples:
There should be no ambiguity between a bare function name and a literal.
Whitespace
Accept insignificant surrounding whitespace, including:
Whitespace inside quoted literals remains significant. The space separator inside a datetime literal is part of the literal format.
Breaking changes from current Expressif syntax
The existing Expressif parser accepts unquoted textual values such as:
This grammar deliberately removes that ambiguity. Equivalent syntax becomes:
Boolean literals are explicitly written as
#trueand#false. Temporal literals are explicitly written using#"...".These changes ensure that bare alphabetic tokens have one syntactic meaning: function calls.
Out of scope
Do not implement or special-case:
@foo)[foo])#1)name := ...)...).foo)$0)|>)Tests
Add Tree-sitter corpus tests covering at minimum:
#trueand#false;trueandfalseparsed as function calls;Tand space;Tests should assert syntax-tree structure, not only parse success.