All posts
RAG Engineering· 6 min read

Hybrid RAG: Why Vector Search Alone Isn't Enough

Hybrid RAG combines semantic (vector) search and keyword search so each covers the other's blind spots. Here we cover the concept, its benefits, what to watch for in Korean-language services, and how to decide whether you need it.

Jiheung Jeong, CEO of SJ System
By Jiheung Jeong
CEO, SJ System · Backend & AI developer, 16 years

One of the most common ways to keep an AI chatbot from confidently making things up is RAG (Retrieval-Augmented Generation): before answering, it looks up evidence in your own material and answers from that evidence. In production, though, answer quality often depends less on the language model and more on retrieval. If the right evidence isn't found, even the best model either says it doesn't know or answers from the wrong source.

This post explains hybrid RAG — a widely used way to improve retrieval quality — what it is, why it matters, and what Korean-language services in particular need to watch for. SJ System also uses a hybrid approach in STAYTOQ, our AI chatbot for lodging businesses.

What RAG is — retrieval sets the ceiling

RAG (Retrieval-Augmented Generation) retrieves relevant material and hands it to the model before it answers. The flow is simple:

  1. Split your material (FAQs, guides, policies) into searchable pieces and store them
  2. When a question arrives, search for the relevant pieces
  3. The model writes the answer based on what was found

A model can only be as accurate as the evidence it receives. So improving RAG mostly means improving step 2 — retrieval.

Why vector search alone falls short

The default in RAG today is vector (dense) search: an embedding model turns text into numeric vectors, and we look for material close in meaning to the question. Its biggest strength is finding matches even when the words differ — "Can I bring my pup?" and "pet policy" share no words but land close together.

Real questions, however, often hinge on exact wording: product names, prices, headcounts, times and proper nouns. Vectors capture the overall meaning of a sentence but tend to be insensitive to such details, so two questions that differ by a single number can look almost the same.

Keyword (sparse) search — scoring by how many words overlap — is the opposite: strong on exact wording, but it misses the same meaning expressed in different words. Their strengths and weaknesses are mirror images.

Vector search (dense)Keyword search (sparse)
Matches onCloseness of meaningOverlapping words
Strong atReworded questions, casual phrasing, slangProduct names, prices, numbers, proper nouns
Weak atQuestions where exact wording mattersSame meaning in different words
Example"Can I bring my doggo?" → pet policy"Rate for building A rooms" → that room's price table

Hybrid search uses both and merges the results. Keyword search covers vector search's blind spots, and vice versa.

How to merge the two

The two searches score in different ways and on different scales. Adding the scores directly lets one side easily overwhelm the other, and keeping the balance right takes constant tuning.

A common alternative is to merge by rank instead of score. A well-known method is RRF (Reciprocal Rank Fusion), which gives more points to items that rank higher in each search and adds them up. Items ranked high in both rise to the top, and items found by only one search can still stay in the running.

Exactly how — and with what weighting — to merge depends on your material and the kinds of questions you get. There's no single right answer, so it's important to decide by measuring against your own real questions.

What Korean-language services need to watch for

In Korean, particles attach directly to nouns. The word for 'dog' with no particle, with an object particle, or with 'also' is the same word to a person, but three different words to a search engine that only splits on spaces. Bring over keyword search built for English as-is, and you can easily end up calling it hybrid while effectively running vector search alone.

So Korean services need keyword search that recognizes words with particles or endings attached as the same word. There are several ways to do this, each with trade-offs, so the choice should fit your material and questions. We ran into keyword search not working as expected early on too, and retrieval quality improved noticeably once we fixed it.

Users don't ask in the words your material uses

Real questions are full of abbreviations, slang, typos and references to earlier messages ("how much is it there?"). A step that tidies the question into a search-friendly form before searching can change results dramatically with the same search engine. Every industry has its own expressions, so this part needs steady improvement based on real usage.

How a hybrid RAG pipeline is put together

Implementations vary, but hybrid RAG generally follows these steps:

  1. Tidy the question — into a search-friendly form
  2. Two kinds of search — by meaning (vector) and by wording (keyword)
  3. Merge the results — into a single ranking
  4. (Optional) Re-rank — refine the top candidates once more
  5. Generate the answer — from the selected evidence

What you use at each step and how you combine them ultimately decides quality, cost and response time. How you split your material into stored pieces (chunking) matters just as much, because the same material can surface different evidence depending on how it's split.

Verify with numbers

Don't judge a retrieval change by whether it "feels better". Prepare questions real users would ask and compare, before and after the change, how often the correct source lands within the top k results (recall@k).

StageShare with the correct source in the top OO
Before improving Korean keyword searchOO%
After improving Korean keyword searchOO%
After refining result rankingOO%
STAYTOQ retrieval quality over time (questions reworded the way guests talk)

Measuring also shows where the problem is. If the right answer isn't among the candidates at all, the issue may be missing material rather than the search method — and adding material is often faster than changing retrieval. Before adding expensive components, we recommend measuring whether you really need them.

What running it taught us

  • Don't automatically block answers when the retrieval score is low. A low score often means a differently worded question, not wrong evidence. Blocking alone throws away correct answers too.
  • Handle Korean from the start. English-oriented defaults leave keyword search unable to do its job.
  • Keep improving question tidying. Industry expressions and slang keep growing as you operate.
  • Measure every retrieval change. Improvements proven in numbers make the next decision easier.

Does your service need hybrid RAG?

If any of the following apply, vector search alone will likely miss questions:

  • Your material has lots of exact terms — product or model names, codes, prices, dates
  • Users ask about the same thing in many different words or slang
  • You search in a language where particles or endings attach to words, such as Korean
  • The work must not be made up — reservations, refunds, policies

SJ System applies hybrid RAG in STAYTOQ and runs it every day at real lodging businesses. If you need an AI that answers precisely from your own material — internal document search or a customer support chatbot — we can start by diagnosing where RAG fits in your work.

  • #RAG
  • #Hybrid search
  • #Vector search
  • #Keyword search
  • #Korean search

Need an AI that answers precisely from your own material?

We start by diagnosing where RAG fits in your work — and the team that builds it runs it to the end.