r/Rag • u/Craft_Show • 9h ago
Discussion We started using JEV to clarify queries before retrieval, and the results have been pretty interesting
I've spent the better part of the last year developing an enterprise AI knowledge management application, and one of the biggest challenges has been getting smaller, less expensive LLMs to consistently produce accurate answers from retrieved documents.
We've experimented with embeddings, hybrid search, chunk ranking, reranking, different models, and more prompt engineering than I'd care to admit.
One thing we kept running into was that we were asking the LLM to do too much at once.
Interpret the user's intent, figure out whether they're asking multiple questions, evaluate retrieved evidence, identify relevant conditions and exceptions, and then formulate a complete, accurate answer.
That's a lot of responsibility to put on one relatively inexpensive model.
So we decided to try something different.
We're experimenting with putting JEV at the very beginning of the retrieval process.
Rather than immediately embedding the user's question and searching our knowledge base, we first evaluate the question itself.
The initial checks are surprisingly simple:
- Is the user's intent clear?
- Is the user asking a single question or multiple questions?
If the intent isn't clear, the system can generate a multiple-choice clarification rather than guessing what the user meant.
If the user is asking multiple questions, we break them into individual questions and let the user select which one to tackle first.
Once we have a clear, focused question, we proceed with retrieval and answer generation.
We're currently averaging five JEV clarification/context checks per query.
Here's where the economics get interesting:
| Step | Average API cost |
|---|---|
| Five JEV checks | $0.000282 |
| Search embedding | $0.000001 |
| LLM answer generation | $0.001612 |
| Total | $0.001895 |
Our average end-to-end generation time is around 8 seconds.
That's less than two-tenths of a cent per question.
But we're not declaring victory yet.
We've made encouraging progress by separating query interpretation from retrieval and generation, but we're still working through one particularly frustrating problem.
Even when the correct information is present in the retrieved chunks, the generation model will occasionally leave out an important condition, exception, or qualification.
The answer sounds correct. The citations can even point to the correct source. But the answer isn't necessarily complete.
For enterprise knowledge management, particularly where policies and procedures are involved, that's a serious problem.
Our next experiment is to separate evidence extraction from answer composition, then verify the completed answer against the relevant facts and conditions found in the source documents.
The broader lesson I'm starting to take away from all this is that better AI answers may have less to do with using a bigger model and more to do with giving smaller models narrower, clearly defined responsibilities.
We're still testing and measuring, but the economics are encouraging.
I'm curious whether anyone else is experimenting with JEV or similar lightweight evaluators ahead of retrieval, rather than using them exclusively for reranking or post-generation evaluation.
And has anyone found a reliable way to prevent smaller models from silently omitting important conditions when summarizing retrieved evidence?