FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing?
Weimin Fu ⋅ Hejia Zhang ⋅ Minghao Shao ⋅ Zeng Wang ⋅ Johann Knechtel ⋅ Ozgur Sinanoglu ⋅ Muhammad Shafique ⋅ Ramesh Karri ⋅ Xiaolong Guo
Abstract
Can large language models generate not just correct, but fast hardware? This paper investigates the question in the context of financial FPGA design, where 5--10 nanoseconds of latency difference determines competitive advantage and designs undergo continuous iteration as exchange protocols, trading strategies, and regulatory requirements evolve. FinHardBench, a benchmark of 33 financial computing tasks, is presented together with three experiments that mirror the real-world FPGA iteration cycle: generating new modules from specifications, tuning system-level configurations across a 6-stage trading pipeline, and adapting existing modules to specification changes. Evaluation of six LLMs on 1530+ experiment rounds yields three findings: (1) models achieve 19--61\% functional correctness with timing degradation up to 13.7$\times$ on specific tasks; (2) in system-level design space exploration, top LLMs converge to the optimal configuration with higher reliability than random search, simulated annealing, and Bayesian optimization baselines (5/5 seeds vs.\ 1--3/5); (3) strategy-level specification changes remain unsolved for most models. Code generation and architectural decision-making are partially independent ($\rho=-0.37$), with task difficulty predicted by training data pattern availability rather than abstraction level. FinHardBench is released as an open-source benchmark.
Successful Page Load