SALA: Syntax-Aware Logit Adjustment for Open-Ended Text Generation
Abstract
Large language models (LLMs) achieve strong performance in text generation but exhibit pronounced syntactic bias, relying heavily on constructions frequently observed during pre-training while underproducing many patterns common in human-written text. Existing decoding strategies primarily operate at the token level through probability truncation and do not explicitly account for linguistic structure, partly due to the computational cost of evaluating syntactic properties during decoding. We propose SALA, a syntax-aware decoding framework that adjusts token logits to encourage the generation of syntactically rare constructions. SALA employs a lightweight linear probe that predicts part-of-speech information from LLM hidden states and token embeddings, enabling efficient syntax-aware re-weighting of token probabilities with minimal computational overhead. The adjustment magnitude is further modulated by the model’s confidence to preserve fluency and coherence. Experiments across multiple models, decoding strategies, and benchmarks demonstrate that SALA consistently improves generation quality and diversity, and yields gains on reasoning tasks, while incurring minimal latency. Our results highlight the importance of incorporating linguistic structure into decoding and suggest that lightweight, post-hoc interventions can effectively mitigate syntactic bias in LLMs.