Advanced RAG chatbot - from vector search to context validation.

Developing an AI chatbot that provides local government information. Designing a pipeline for advanced development.

·~6 min read·687 views·Updated: Oct 5, 2026

I created an AI chatbot that provides local government information. It started as a simple RAG (Retrieval-Augmented Generation) pipeline, but while testing actual questions, I discovered several issues and enhanced it into a 7-step pipeline to resolve them.

Limitations of Basic RAG

Having previously worked on a project using an RAG pipeline and orchestration with AWS Bedrock, I wanted to build the RAG pipeline myself this time. The initially constructed RAG pipeline was simple.

User query -> Embedding -> Vector search -> LLM answer generation
User query -> Embedding -> Vector search -> LLM answer generation

However, the above method had several problems.

Failure to Understand Context

User: "I don't have a washing machine at home, is there a laundry service?"
Chatbot: "In XX district, there is an appliance support program for low-income households..."
User: "I don't have a washing machine at home, is there a laundry service?"
Chatbot: "In XX district, there is an appliance support program for low-income households..."

In the actual prepared data, there was information related to 'support for laundry and luggage storage services,' but it got stuck on the keyword washing machine and generated an irrelevant response. (In fact, there was no such support program.)

Ignoring Conversation History

User: "What welfare benefits are available for the disabled?"
Chatbot: (Provides information on welfare for the disabled)

User: "Until when can I apply??"
Chatbot: "I'm sorry. I cannot find information on the application deadline."
User: "What welfare benefits are available for the disabled?"
Chatbot: (Provides information on welfare for the disabled)

User: "Until when can I apply??"
Chatbot: "I'm sorry. I cannot find information on the application deadline."

Since the vector search was done based on the embedded user query, the context resulting from the vector search for the follow-up question itself was sometimes unrelated to or did not fit the previous question history.

7-Step Pipeline Design

To solve the above problems, I redesigned the pipeline as follows.

[Step 1] Receive user query
↓
[Step 2] Reconstruct query based on conversation history (using LLM)
↓
[Step 3] Generate embedding
↓
[Step 4] Vector search (top_k = 5)
↓
[Step 5] Combine all chunks of related documents
↓
[Step 6] Include attachment metadata
↓
[Step 7] Validate LLM context and generate answer
[Step 1] Receive user query
↓
[Step 2] Reconstruct query based on conversation history (using LLM)
↓
[Step 3] Generate embedding
↓
[Step 4] Vector search (top_k = 5)
↓
[Step 5] Combine all chunks of related documents
↓
[Step 6] Include attachment metadata
↓
[Step 7] Validate LLM context and generate answer

Step 2: Reconstruct Query Based on Conversation History

Based on the conversation with the chatbot, I aimed to reconstruct the incoming user query to fit the context.

The results were as follows.

User: "What welfare is available for the disabled?"
→ Reconstructed query: "Types of welfare policies for the disabled"

User: "Until when can I apply?"
→ Reconstructed query: "Application period for welfare policies for the disabled"  # Maintains context based on previous question!!
User: "What welfare is available for the disabled?"
→ Reconstructed query: "Types of welfare policies for the disabled"

User: "Until when can I apply?"
→ Reconstructed query: "Application period for welfare policies for the disabled"  # Maintains context based on previous question!!

Step 5: Combine All Chunks of Related Documents

I believe this was the core of the project. Previously, only the chunks resulting from the vector search were used as context, which led to insufficient data to reference when generating answers.
The data stored in the vector DB was chunked into fixed-size chunks of about 500 tokens with a 30-token overlap, but I determined that providing accurate information with just those chunks was difficult, so I decided to combine all chunks from the original document to which the searched chunk belonged to provide as context.

[Step 4] Vector search completed - Number of results: 5
  [1] Similarity: 0.4939 | Title: Support for laundry and luggage storage... | ID: doc_123
  [2] Similarity: 0.4062 | Title: Support for home care services... | ID: doc_456
[Step 5] Combine all chunks of related documents
  - doc_123: Combined 3 chunks
  - doc_456: Combined 2 chunks
[Step 4] Vector search completed - Number of results: 5
  [1] Similarity: 0.4939 | Title: Support for laundry and luggage storage... | ID: doc_123
  [2] Similarity: 0.4062 | Title: Support for home care services... | ID: doc_456
[Step 5] Combine all chunks of related documents
  - doc_123: Combined 3 chunks
  - doc_456: Combined 2 chunks

To provide accurate answers, I concluded that rather than generating answers from just a part of the document, it was necessary to find similar relationships through vector search and generate answers based on the entire content of the document.

Step 7: LLM Context Validation

Before generating the actual answer to be provided to the user, I used LLM to verify whether the vector search results were indeed related to the user query and whether the data was appropriate before providing the answer.

Received a response in the following format through LLM.
- is_relevant: true/false  // Whether the context is appropriate
- confidence: high/medium/low // Confidence level
- action: answer/refine_search/no_data // Proceed with answer, re-search, no data found notification
Received a response in the following format through LLM.
- is_relevant: true/false  // Whether the context is appropriate
- confidence: high/medium/low // Confidence level
- action: answer/refine_search/no_data // Proceed with answer, re-search, no data found notification

Technology Stack

  • Python FastAPI

  • FAISS (vector store)

  • jhgan/ko-sroberta-multitask (Korean specialized embedding open source)

  • React

Final Log Example

INFO: RAG streaming query: Tell me about disability benefits (session: ea2431fb...)
INFO:=======================================================================
INFO: [Step 1] User query: Tell me about disability benefits
INFO: [Step 2] Reconstructed search query: Content and application method for disability benefits
INFO: [Step 3] Embedding generation completed
INFO: [Step 4] Vector search completed - Number of results: 5
INFO:   [1] Similarity: 0.6234 | Title: Support for disability benefits... | Department: Disability Welfare Division
INFO:   [2] Similarity: 0.5891 | Title: Guide to disability pensions... | Department: Disability Welfare Division
INFO: [Step 5] Combine all chunks of related documents - 2 documents, 7 chunks
INFO: [Step 6] Check attachments - 1 document included
INFO: [Step 7] LLM validation completed
INFO:   - Relevance: True
INFO:   - Confidence: high
INFO:   - Action: answer
INFO: Starting answer generation...
INFO: RAG streaming query: Tell me about disability benefits (session: ea2431fb...)
INFO:=======================================================================
INFO: [Step 1] User query: Tell me about disability benefits
INFO: [Step 2] Reconstructed search query: Content and application method for disability benefits
INFO: [Step 3] Embedding generation completed
INFO: [Step 4] Vector search completed - Number of results: 5
INFO:   [1] Similarity: 0.6234 | Title: Support for disability benefits... | Department: Disability Welfare Division
INFO:   [2] Similarity: 0.5891 | Title: Guide to disability pensions... | Department: Disability Welfare Division
INFO: [Step 5] Combine all chunks of related documents - 2 documents, 7 chunks
INFO: [Step 6] Check attachments - 1 document included
INFO: [Step 7] LLM validation completed
INFO:   - Relevance: True
INFO:   - Confidence: high
INFO:   - Action: answer
INFO: Starting answer generation...

In Conclusion

Starting from a simple RAG project and contemplating an advanced orchestration processing method like AWS Bedrock's Agent service, I went through the process of enhancing it into a 7-step pipeline, and I noticed a significant improvement in responses with each additional step.

Share