The 2nd International Workshop on Language Models and Programming Languages (LMPL 2026), Oakland, USA
by Lara Marinov1, Aditya Thimmaiah1, Jayanth Srinivasa2, Junyi Jessy Li1, Milos Gligoric1
1The University of Texas at Austin 2Cisco Research
Do LLMs actually understand programming language semantics when completing software engineering tasks or do they just fall back on patterns they've memorized in training?
To answer this, we introduce Program Executability Prediction (PrEx): given a program’s syntax and operational semantics, decide whether the program is semantically valid or invalid—and if invalid, which formal rule it violates.
The details for requirements and setup are contained in the PLSemanticsBench repository:
Basic usage to rerun experiments:
python src/main.py experiment <experiment_name>
Example:
python src/main.py experiment qwen_coder_32b_uk_pcp_imp_k
For CoT/multi-seed runs, add --seed:
python src/main.py experiment qwen_coder_3b_mk_pcp_cot_imp_sos --seed 73
<experiment_name> is any registered subcommand in src/llm_interpreter/experiments/functions.py (e.g. qwen_coder_*_pcp_*, ministral_*_pcp_*, deepseek_*_pcp_*).
@inproceedings{MarinovETAL26PrEx,
author = {Marinov, Lara and Thimmaiah, Aditya and Srinivasa, Jayanth and Li, Junyi Jessy and Gligoric, Milos},
title = {Predicting Program Exit Code with {LLMs} and Programming Language Semantics},
booktitle = {International Workshop on Language Models and Programming Languages},
pages = {To appear},
year = {2026},
}