Marcio Cunha

Mitigating Hallucinations in Language Models with Sparse and Dense Vector Retrieval

Learn how to combine exact keyword-based sparse searches and deep semantic-based dense searches to eliminate artificial intelligence hallucinations.

Marcio Cunha•5 min
Also available in:PortuguêsEspañol
Summary
  • Hybrid retrieval unites exact keyword matching for rare terms with deep comprehension of semantic context.
  • Language models invent facts due to a lack of precise data in the knowledge base or lost focus in the prompt.
  • Dense vectors find related ideas, while sparse vectors guarantee exact codes, acronyms, and proprietary terms.
  • Weighted reciprocal rank fusion balances the weight of both approaches before sending context to the model.
  • Implementing this architecture drastically reduces incorrect answers without requiring full retraining of neural networks.

The Challenge of Hallucinations in Artificial Intelligence Systems

When interacting with a modern language model, textual fluency often masks a critical flaw: the invention of facts, known in technical terms as a hallucination. In practice, this means the artificial intelligence generates responses that look perfectly correct but are entirely false. This behavior occurs because models are trained to predict the most likely next word based on statistical patterns rather than consulting a verifiable database at runtime. To mitigate this issue without altering the internal neural network structure, software engineering adopted Retrieval-Augmented Generation, which involves searching for relevant external documents before generating the response.

However, traditional document retrieval faces severe technical barriers when dealing with highly specific terms such as product codes, corporate acronyms, or rare proper nouns. If the system searches only for the general meaning of words, it may miss crucial documents containing the exact term typed by the user. Conversely, searches strictly based on exact keyword matching fail miserably when a user expresses a question using synonyms or sentence structures completely different from company records. It is precisely within this crossed limitation scenario that combining complementary data indexing approaches becomes necessary to feed the language model with clean, precise context.

Understanding Word-Based Sparse Retrieval

Sparse retrieval is the classic search approach, frequently associated with traditional algorithms like BM25. In practice, it works much like an index at the back of a technical book: the technology analyzes text, counts word frequencies, and builds an inverted list mapping which documents contain which terms. It is called sparse because the resulting mathematical matrix contains thousands of columns corresponding to the total vocabulary, while most values remain zero, indicating the term is absent in that specific document. This technique is unbeatable for finding exact terms, serial numbers, error codes, and function names within a massive repository of technical documentation.

Despite high lexical precision, sparse retrieval suffers from deep semantic myopia. If technical documentation uses timeout error and the user types system frozen, the traditional sparse algorithm may ignore the correct document due to a lack of exact word overlap, even if the meaning is identical. Furthermore, grammatical variations and plurals require additional normalization rules to prevent search failures. For this structural reason, relying solely on sparse searches leaves dangerous gaps that allow language models to invent data to fill the informational void delivered in the context.

The Revolution of Dense Vectors and Semantic Search

In direct contrast to word-counting systems, dense retrieval uses embedding models to translate the deep meaning of sentences and paragraphs into numeric sequences of hundreds or thousands of dimensions. In practice, think of these vectors as spatial coordinates where concepts with similar meanings are physically close, regardless of the exact vocabulary used. If a document discusses automobile and the query mentions car, the dense model perceives that the points in the vector space are very close and retrieves the content with high fidelity. This capability to capture the intention behind writing transforms how systems find relevant information in unstructured databases.

However, pure dense search has operational vulnerabilities that can sabotage corporate assistant reliability. Because vectors compress the global meaning of text into generalized numeric coordinates, they tend to lose granular details such as exact part codes, specific software library versions, or tax identification numbers. If a developer asks about compilation error ERR-4091, a dense model may return documents about compilation errors in general while omitting the exact document covering that specific code. This loss of surgical precision forces the language model to guess the remaining information, directly generating the hallucination we wanted to avoid.

The Hybrid Architecture: Uniting the Best of Both Worlds

To simultaneously solve semantic myopia and exact term loss, engineers adopted hybrid retrieval, which executes sparse and dense searches in parallel. In practice, the system receives the user query, sends one route to the traditional word-based search engine and another to the semantic-based vector database. Each engine returns its own list of candidate documents accompanied by distinct relevance scores. The secret of this architecture lies in the fusion stage, where these heterogeneous scores are normalized and combined into a single unified list, ensuring both broad conceptual context and exact engineering terms appear at the top of the results.

The most robust technique for this fusion is Weighted Reciprocal Rank Fusion, an algorithm that reorganizes lists based on the positions documents achieved in each individual search. If a document ranked first in the sparse search due to the exact code and second in the dense search due to general context, it receives a very high combined score and moves straight to the top of the context pile. This flow eliminates the blind spots of each isolated technology, providing the language model with highly relevant, factual, and verifiable context snippets, cutting off the need for the system to invent answers at the root.

Implementing Result Fusion in Python
To demonstrate how search fusion works in code, let us implement a simple reciprocal score function combining sparse and dense search results. In practice, this function takes document ID lists returned by each engine and calculates a consolidated score to order the final documents.

def hybrid_fusion(sparse_results, dense_results, k=60, weight_sparse=0.5, weight_dense=0.5):
fusion_scores = {}

for rank, doc_id in enumerate(sparse_results):
score = weight_sparse * (1.0 / (k + rank + 1))
fusion_scores[doc_id] = fusion_scores.get(doc_id, 0.0) + score

for rank, doc_id in enumerate(dense_results):
score = weight_dense * (1.0 / (k + rank + 1))
fusion_scores[doc_id] = fusion_scores.get(doc_id, 0.0) + score

sorted_docs = sorted(fusion_scores.items(), key=lambda x: x[1], reverse=True)
return [doc_id for doc_id, score in sorted_docs]

This code snippet illustrates the core score combination algorithm without requiring complex database dependencies to understand the mathematical logic. By adjusting weights, the engineer can prioritize exact term matching or semantic proximity as application domains require higher technical rigor.

Final Considerations on Reliability and Scale

Eliminating hallucinations in artificial intelligence systems is no longer purely an academic research problem but a fundamental software engineering requirement for production applications. By abandoning the illusion that a single data retrieval method can meet all nuances of human language, technology teams gain robustness by adopting hybrid sparse and dense vector architectures. This synergy ensures that both abstract concepts and specific technical details reach the model intact, establishing a solid foundation of verifiable facts for secure, precise answers.

Maintaining this infrastructure at scale requires constant monitoring of retrieved context quality and fine-tuning fusion parameters as the corporate knowledge base grows. With a clean, well-indexed data foundation retrieved with surgical precision, language models fulfill their role predictably and reliably, delivering real value to users without the unwanted risks of invented responses.