Hybrid RAG: Why Vector Search Alone Isn't Enough
Hybrid RAG combines semantic (vector) search and keyword search so each covers the other's blind spots. Here we cover the concept, its benefits, what to watch for in Korean-language services, and how to decide whether you need it.

One of the most common ways to keep an AI chatbot from confidently making things up is RAG (Retrieval-Augmented Generation): before answering, it looks up evidence in your own material and answers from that evidence. In production, though, answer quality often depends less on the language model and more on retrieval. If the right evidence isn't found, even the best model either says it doesn't know or answers from the wrong source.
This post explains hybrid RAG — a widely used way to improve retrieval quality — what it is, why it matters, and what Korean-language services in particular need to watch for. SJ System also uses a hybrid approach in STAYTOQ, our AI chatbot for lodging businesses.
What RAG is — retrieval sets the ceiling
RAG (Retrieval-Augmented Generation) retrieves relevant material and hands it to the model before it answers. The flow is simple:
- Split your material (FAQs, guides, policies) into searchable pieces and store them
- When a question arrives, search for the relevant pieces
- The model writes the answer based on what was found
A model can only be as accurate as the evidence it receives. So improving RAG mostly means improving step 2 — retrieval.
Why vector search alone falls short
The default in RAG today is vector (dense) search: an embedding model turns text into numeric vectors, and we look for material close in meaning to the question. Its biggest strength is finding matches even when the words differ — "Can I bring my pup?" and "pet policy" share no words but land close together.
Real questions, however, often hinge on exact wording: product names, prices, headcounts, times and proper nouns. Vectors capture the overall meaning of a sentence but tend to be insensitive to such details, so two questions that differ by a single number can look almost the same.
Keyword (sparse) search — scoring by how many words overlap — is the opposite: strong on exact wording, but it misses the same meaning expressed in different words. Their strengths and weaknesses are mirror images.
| Vector search (dense) | Keyword search (sparse) | |
|---|---|---|
| Matches on | Closeness of meaning | Overlapping words |
| Strong at | Reworded questions, casual phrasing, slang | Product names, prices, numbers, proper nouns |
| Weak at | Questions where exact wording matters | Same meaning in different words |
| Example | "Can I bring my doggo?" → pet policy | "Rate for building A rooms" → that room's price table |
Hybrid search uses both and merges the results. Keyword search covers vector search's blind spots, and vice versa.
How to merge the two
The two searches score in different ways and on different scales. Adding the scores directly lets one side easily overwhelm the other, and keeping the balance right takes constant tuning.
A common alternative is to merge by rank instead of score. A well-known method is RRF (Reciprocal Rank Fusion), which gives more points to items that rank higher in each search and adds them up. Items ranked high in both rise to the top, and items found by only one search can still stay in the running.
Exactly how — and with what weighting — to merge depends on your material and the kinds of questions you get. There's no single right answer, so it's important to decide by measuring against your own real questions.
What Korean-language services need to watch for
In Korean, particles attach directly to nouns. The word for 'dog' with no particle, with an object particle, or with 'also' is the same word to a person, but three different words to a search engine that only splits on spaces. Bring over keyword search built for English as-is, and you can easily end up calling it hybrid while effectively running vector search alone.
So Korean services need keyword search that recognizes words with particles or endings attached as the same word. There are several ways to do this, each with trade-offs, so the choice should fit your material and questions. We ran into keyword search not working as expected early on too, and retrieval quality improved noticeably once we fixed it.
Users don't ask in the words your material uses
Real questions are full of abbreviations, slang, typos and references to earlier messages ("how much is it there?"). A step that tidies the question into a search-friendly form before searching can change results dramatically with the same search engine. Every industry has its own expressions, so this part needs steady improvement based on real usage.
How a hybrid RAG pipeline is put together
Implementations vary, but hybrid RAG generally follows these steps:
- Tidy the question — into a search-friendly form
- Two kinds of search — by meaning (vector) and by wording (keyword)
- Merge the results — into a single ranking
- (Optional) Re-rank — refine the top candidates once more
- Generate the answer — from the selected evidence
What you use at each step and how you combine them ultimately decides quality, cost and response time. How you split your material into stored pieces (chunking) matters just as much, because the same material can surface different evidence depending on how it's split.
Verify with numbers
Don't judge a retrieval change by whether it "feels better". Prepare questions real users would ask and compare, before and after the change, how often the correct source lands within the top k results (recall@k).
| Stage | Share with the correct source in the top OO |
|---|---|
| Before improving Korean keyword search | OO% |
| After improving Korean keyword search | OO% |
| After refining result ranking | OO% |
Measuring also shows where the problem is. If the right answer isn't among the candidates at all, the issue may be missing material rather than the search method — and adding material is often faster than changing retrieval. Before adding expensive components, we recommend measuring whether you really need them.
What running it taught us
- Don't automatically block answers when the retrieval score is low. A low score often means a differently worded question, not wrong evidence. Blocking alone throws away correct answers too.
- Handle Korean from the start. English-oriented defaults leave keyword search unable to do its job.
- Keep improving question tidying. Industry expressions and slang keep growing as you operate.
- Measure every retrieval change. Improvements proven in numbers make the next decision easier.
Does your service need hybrid RAG?
If any of the following apply, vector search alone will likely miss questions:
- Your material has lots of exact terms — product or model names, codes, prices, dates
- Users ask about the same thing in many different words or slang
- You search in a language where particles or endings attach to words, such as Korean
- The work must not be made up — reservations, refunds, policies
SJ System applies hybrid RAG in STAYTOQ and runs it every day at real lodging businesses. If you need an AI that answers precisely from your own material — internal document search or a customer support chatbot — we can start by diagnosing where RAG fits in your work.
- #RAG
- #Hybrid search
- #Vector search
- #Keyword search
- #Korean search
Need an AI that answers precisely from your own material?
We start by diagnosing where RAG fits in your work — and the team that builds it runs it to the end.