But there was a catch. Latency was prohibitive, and so was the token cost. So, we landed on a hybrid approach. We used the LLM to generate labeled data to train an ML classifier. We got the best of both worlds, solving the cold-start problem while avoiding the latency and costs of large language models. That’s the system that ultimately won in production. The LLM earned its place at build time, not at request time.

The continuum, and how to choose

It is tempting to take a powerful solution, such as LLMs and AI agents, and simply apply it to every problem. But when we look at the four use cases above, each of them led to a different decision. In the recommender system, GenAI did not earn its place. In content moderation, it earned part of the job, providing the explanation while the ML model provided the score. In querying the database, the agent did the reasoning, while the semantic layer pulled the right number. And in search, the LLM solved the cold-start problem, but it was never deployed to production.

Behind all of these decisions is the same underlying question: where does this technology earn its place, and does it justify the added complexity?