Stanford AI Lab recently highlighted a new paper called "The Embedder's Dilemma", from Adnan El Assadi, Niklas Muennighoff and Jinhyuk Lee, asking a deceptively simple question: now that LLMs can perform many of the tasks we traditionally hand to embedding models, should we just use the LLM?
Their answer is more interesting than yes or no.
Across 37 tasks, the best LLM and the best embedding model were effectively tied overall. But the averages hide an important split: embedding models remained stronger for classification, the two approaches were comparable across several similarity tasks, while LLMs pulled ahead on reasoning-heavy retrieval.
The catch? Getting there can cost enormously more.
In the researchers' evaluation, an LLM could cost up to 1,431 times more than a comparable-quality embedding model. The open LLMs they tested were also between 2.5 and 736 times slower than embedding models when evaluated on the same GPU.
That crystallises a design decision engineering teams are increasingly facing: when does a query actually deserve the expensive model, and when does the cheap one already do the job?
This isn't a philosophical question. It shows up in your cloud bill, your p95 latency, and your users' patience.
An embedding model takes a piece of text, turns it into a vector of numbers, and allows you to find nearby vectors in an index. It's fast, cheap, and stateless. You run it once per document at index time, then again at query time, and relatively inexpensive nearest-neighbour arithmetic does the rest.
An LLM can approach retrieval differently. Instead of reducing every document and query to fixed vectors and comparing their proximity, it can inspect text directly and use its language and reasoning capabilities to judge relevance.
That distinction becomes important when relevance itself requires reasoning.
"Find documents mentioning cancellation" is largely a similarity problem.
"Which of these documents would help someone determine whether they can cancel after their renewal date?" requires understanding relationships that may never be expressed using similar words.
And there is another distinction worth making. An LLM can also operate after retrieval: interpreting the documents that were found, re-ranking them, combining information from several sources, or generating an answer for the user. That is a different job from retrieval itself, even though modern AI applications often bundle the two together.
The price of all that reasoning is compute.
The Stanford paper doesn't say embeddings have suddenly become obsolete. In fact, its results suggest almost the opposite: the two approaches have different strengths, and the expensive model earns its price only when the task actually benefits from what it can do.
Pull up your query logs and you'll probably find that many of them fall into two broad groups. A large proportion are short, specific, and well-formed: "reset password," "invoice #4471," "cancel subscription," "Python string split." For these, an embedding model is not a compromise, it's the correct tool.
The characteristics that make embeddings the right call:
In these conditions, a capable embedding model handles retrieval cleanly. The "cheaper" option isn't a sacrifice; it's a sensible match between the tool and the task.
And cheap is not a minor distinction here. In the paper's evaluation, the open LLMs tested were anywhere from 2.5 to 736 times slower than embedding models on the same GPU. At production scale, even a modest quality improvement has to justify that difference in compute, latency and cost.

The other cluster in those logs looks quite different.
Consider queries like "which of our pricing plans makes sense for a ten-person team that needs SSO but doesn't need analytics," or "what's the difference between feature A and feature B for someone coming from Competitor X."
These are not simple similarity lookups. Determining which documents are relevant may itself require reasoning about what the user means.
And then there are questions like "summarise what changed in our product between Q1 and Q3." Here, finding the relevant documents is only half the job. The user's actual request is synthesis.
This is an important distinction because an LLM can earn its place in two different parts of the pipeline.
It can improve retrieval itself when determining relevance requires reasoning. Or it can operate after retrieval, reasoning over the documents that have already been found and producing the answer the user actually wants.

An LLM call becomes worth considering when:
There is also a useful middle ground: retrieve cheaply with embeddings first, then let an LLM re-rank the smaller candidate set. The paper evaluates this retrieve-then-rerank approach as well, and it illustrates an important architectural principle. You don't necessarily have to choose between embeddings and an LLM for the entire pipeline.
I find this genuinely clarifying: the question is not "which model is better" in the abstract. It's "what is the query actually asking for, and what does a good answer look like?"
The shape of the right answer determines the tool.

Once you stop treating this as an either-or decision, the architecture becomes more interesting.
Straightforward retrieval can go to the embedding index. Queries where relevance itself requires reasoning can take a more expensive path. Some queries can use embeddings to retrieve a candidate set and an LLM to re-rank it. And questions that require an actual synthesised answer can retrieve first and then pass the relevant context to an LLM for generation.
In other words, you can think of the system as having several gears rather than one engine running at full throttle all the time.
A routing layer can decide which path a query takes.
The router itself does not need to be particularly intelligent. A lightweight classifier trained on your own query logs, a few carefully designed heuristics, or a small inexpensive language model sitting in front of the main pipeline can all work. You're not asking the router to answer the question; you're asking it to recognise the shape of the question.
There is another result buried in the paper that I think deserves more attention.
More reasoning wasn't automatically better.
Reasoning tokens accounted for 28–81% of LLM inference cost in the researchers' evaluation, yet lowering the reasoning budget preserved or actually improved retrieval performance for most of the models they tested.
That's a useful reminder for anyone building AI systems today. The optimisation problem isn't simply "use the smartest model available." Even after you've decided a query deserves an LLM, you still have to ask: how much intelligence are you actually buying for this particular task?
There is another optimisation the paper isn't about but that matters in production: semantic caching.
If you cache LLM responses against semantically similar queries, repeated questions do not necessarily require repeated inference. Common questions in a SaaS product — pricing queries, onboarding help, feature explanations — recur constantly. If the underlying information has not changed and the application can safely reuse the response, serving a validated cached result can be considerably cheaper than calling the model again.
The important part is that caching, routing and model selection are not separate tricks. They are all expressions of the same architectural principle:
don't spend compute where it doesn't improve the outcome.

If you're a founder or a product team deciding how to build search into your application, the practical takeaway is this: don't default to the most capable model because it feels safer. That instinct is expensive.
Start by categorising what your users are actually searching for.
If most of it is lookup — finding records, articles, products, past conversations — an embedding model may get you most of the way there at a fraction of the cost.
If relevance itself requires reasoning, add an LLM where it earns its place.
If the user wants an answer rather than a document, retrieve the information first and let the LLM synthesise it.
And if you can retrieve cheaply and use the expensive model only to re-rank a small candidate set, that may be a better compromise than making the LLM do everything.
I'd argue the teams that get this right aren't the ones who pick the most powerful model. They're the ones who know exactly which queries need that power and route everything else to something faster and cheaper.
My experience with building software has taught me that the cost discipline you build into an architecture at the start is much easier to maintain than cost discipline you try to retrofit later, when users are already complaining about slow search and your cloud bill has doubled.
There is a broader lesson here too.
Capability is not architecture.
The fact that an LLM can perform a task previously handled by an embedding model does not mean an intelligently designed system should replace the embedding model with it.
That distinction is going to matter more as models become increasingly capable. If every new capability automatically becomes another reason to send every request through the biggest model available, AI applications will become extraordinarily good at converting cloud budgets into heat.
Good architecture asks a different question: what is the least expensive system that reliably produces the outcome we need?
The findings from The Embedder's Dilemma are worth taking seriously. LLMs now have a real advantage on reasoning-heavy retrieval, while embedding models remain extraordinarily competitive — and dramatically cheaper — across much of the rest of the workload.
So the design question isn't whether LLMs have become "better than embeddings." That's the wrong abstraction.
The question is whether the performance delta matters for this query, at this scale, under this latency budget, at this cost tolerance.
Often it won't.
When it does, spend the money without hesitation.

LLMs have not made embedding models obsolete. The Embedder's Dilemma found the best LLM and embedding model effectively tied across its overall evaluation, but with very different strengths and costs. Embedding models remain highly effective for similarity-oriented tasks and straightforward retrieval, while LLMs pull ahead when retrieval itself requires reasoning. The practical architecture is therefore not "LLM or embeddings," but using the cheapest tool that reliably handles each query and escalating to an LLM when its additional capability actually changes the result.
No. Across the paper's 37 evaluation tasks, the best LLM and best embedding model were effectively tied overall. LLMs performed better on reasoning-heavy retrieval, while embedding models remained stronger or competitive across several other task categories at dramatically lower cost. For production systems, the findings support using each where its strengths justify the cost.
An LLM becomes more useful when determining relevance requires reasoning rather than simple semantic similarity. Straightforward document lookup, classification and similarity search remain strong use cases for embedding models, while ambiguous or reasoning-intensive retrieval can justify an LLM or an LLM re-ranking stage.
Start with your own query logs. Short, specific lookup queries can often be handled efficiently with embeddings. Conversational questions, queries requiring relationships between several facts, or questions whose relevance depends on reasoning are stronger candidates for an LLM. A lightweight routing layer can classify queries and escalate only those that need the more expensive path.
It can be enormous. In The Embedder's Dilemma, an LLM could cost up to 1,431 times more than a comparable-quality embedding model in the researchers' evaluation, while the tested open LLMs were between 2.5 and 736 times slower on the same GPU. Those figures come from the paper's particular experimental setup rather than being universal production ratios, so actual costs will depend on your models, hardware, token counts and providers.
No. A hybrid architecture can use embeddings for initial retrieval and then send a smaller candidate set to an LLM for re-ranking. For questions requiring a generated answer, the retrieved documents can then become context for an LLM. This lets the expensive model work only where its additional capability is useful.
Semantic caching stores responses so that sufficiently similar future queries can reuse an existing result instead of triggering another LLM inference. It can be useful for applications with repetitive questions, provided the underlying information is still valid and the application can safely reuse the cached response.
Yes. The router itself can be a lightweight classifier, carefully designed heuristics, or a small language model. The difficult part is not necessarily the size of the infrastructure; it is understanding your query patterns well enough to decide which requests genuinely require more expensive reasoning.