QualAlign: Benchmarking Automated Qualitative Coding Against Human Schemas
Abstract
Recent LLM-based systems can produce qualitative codes and themes quickly from raw text, but it remains unclear how closely their outputs match the artifacts that expert qualitative researchers would produce. We introduce QualAlign, a human-alignment benchmark for automated qualitative coding spanning eight public datasets, twenty-one question-conditioned analysis slices, and segment-level expert annotations. Under a common inference setting, we evaluate four recent systems: LLooM, Thematic-LM, HICode, and LOGOS. We use a multi-metric suite that separates lexical similarity, semantic similarity, code-assignment agreement, and LLM-judged similarity. Results reveal consistent cross-metric disagreement: embedding similarity is high, lexical overlap is low, and assignment-level agreement remains weak. Case studies show that systems often merge multiple human atomic codes into broader labels or split broader human codes into multiple granular variants, preserving topical similarity while reducing agreement. They also show that semantic resemblance can mask interpretive reframing. QualAlign provides a reusable benchmark and supports a human-grounded, multi-metric view of automated inductive coding.