Is NLP System-Building a Machine Learning Problem?
Abstract
Everyone nowadays is building and evaluating LLM workflows. It's easy to get something running, but we don't yet have an engineering discipline for zeroing in on the most accurate and cost-effective system. There are many ways to break a given task into subtasks -- each of which may benefit from evaluation and revision. There are many components to try, including prompts, examples, rewards, LLMs, and external tools -- as well as humans in the loop, human annotators before the loop, and human evaluators after the loop. Many of these components are trainable or configurable and have varying costs. Do we really have to search this space by hand? Or could AI guide our use of annotation and computation?
The problem is more general than NLP. For example, medicine develops workflows for diagnosis and treatment. I'll outline a general approach based on active feature acquisition: every human annotation, LLM output, tool call, or medical test result is a random variable that we could pay to observe. Since random variables are often correlated, observing cheap variables can help us predict more expensive variables and evaluate whether they are worth observing as well. (Which step should we try next, and with what prompt? Should we double-check or revise the result? Should we ask a human?) Ultimately, gathering information is a reinforcement learning problem. I'll describe how to train an environment model that predicts distributions over any missing variables, much as BERT predicts distributions over any missing tokens. This model could be extended into a policy model that chooses which missing variable to observe next. Any data gathered by the system can immediately be used to train it further, including on counterfactual trajectories, which exploit the fact that our actions have no side effects (observing variables does not change the values of other variables).
Speaker