Stvya Sharma
BAC Honours Thesis
Probing Symbolic Reasoning in Large Language Models with Perpetually Renewable, Contamination-Free Grammar Benchmarks
By Stvya Sharma
Supervised by: Rahul Gopinath
Abstract
Large language models (LLMs) excel at informal reasoning yet exhibit pronounced brit- tleness in formal rule-following, frequently defaulting to approximate pattern matching rather than systematic execution. This limitation is compounded by a crisis of measurement: as state-of-the-art models saturate existing benchmarks through training data contamination, researchers must perpetually construct new evaluations to distinguish genuine capability from memorisation. This thesis introduces a perpetually renewable benchmark framework that decomposes formal reasoning into three independently renewable, progressively complex dimensions (symbolic constraint satisfaction, contextual anchoring, and semantic grounding) to isolate the architectural and task-specific factors governing performance. Across three experiments utilising procedurally generated, inexhaustible test instances, we demonstrate that the boundary between genuine reasoning and pattern matching is fundamentally task- dependent, shifting predictably as each dimension of difficulty is introduced. Results reveal that contextual anchoring and semantic pressure differentially shape model behaviour, with rule-following strategies adapting conditionally to the specific constraint composition. By eliminating data contamination risks and enabling longitudinal evaluation, the framework dis- tinguishes inherent incapacities from training-distribution artefacts, establishing a principled, sustainable methodology for characterising systematic reasoning capabilities.