Learning Reusable Program Transformations via LLM-Guided Rule Synthesis
Abstract
Optimizing Pandas programs is a challenging problem. Existing techniques prove ineffective when tested on real-world benchmarks. Using LLMs in a per-program optimization methodology can synthesize nontrivial optimizations, but it is expensive due to its low yield. To address this, we introduce RuleCraft, a 3-stage approach that decouples discovery from deployment and connects them via a novel bridge. First, it discovers and validates per-program optimizations (discovery). Second, they are converted into generalised rewrite rules (bridge). Finally, these rules are incorporated into a compiler that can automatically apply them wherever applicable, eliminating repeated reliance on LLMs (deployment). We demonstrate that RuleCraft is the new state-of-the-art (SOTA) Pandas optimization framework on PandasBench, a challenging Pandas benchmark consisting of Python notebooks. Across these notebooks, we achieve a speedup of up to 4.3x over Dias, the previous compiler-based SOTA, and 1914.9x over Modin, the previous systems-based SOTA.