How Malaysian Enterprises Can Use Localized RAG for Bahasa Melayu & Manglish
Published by: Gautham Krishna RSep 30, 2026Blog
The Problem With "One-Size-Fits-All" AI in Malaysia
A customer types a question into a banking chatbot: "Nak tahu macam mana nak dispute transaction yang tak recognized ni, ada charge RM200." The sentence starts in Bahasa Melayu, pivots into English terminology, and carries a Manglish rhythm. A generic AI system trained predominantly on English data might retrieve irrelevant documents. It might respond in stilted formal Malay that sounds nothing like how Malaysians actually communicate. Or worse, it might hallucinate an answer because it never properly understood the query in the first place.
This is the everyday reality for Malaysian enterprises exploring AI-powered knowledge systems. According to the "Unlocking Malaysia's AI Potential 2026" study commissioned by Amazon Web Services and conducted by Strand Partners, 38% of Malaysian businesses now consistently use at least one AI tool, up from 27% in 2025. More than 3.4 million businesses have adopted AI, yet 67% remain focused on basic applications like public chatbots and ready-made tools. Financial services and manufacturing lead adoption at 53% and 50% respectively, but the gap between experimentation and enterprise-scale deployment remains significant.
The next challenge is turning broad adoption into repeatable, organisation-wide capability. For Malaysian enterprises, that means building AI systems that actually understand how Malaysians write, speak, and search.
What RAG Actually Is (And Why It Matters)
Retrieval-Augmented Generation, or RAG, is a technique that grounds AI responses in specific documents rather than relying solely on what a language model memorized during training. Instead of asking a model to "know" everything, you give it a searchable knowledge base and a retrieval system that finds relevant passages before the model generates an answer.
The workflow looks like this:
Local enterprise data > multilingual preprocessing > embeddings/indexing > retrieval > reranking > LLM generation > grounded response > evaluation
Think of it like a research assistant who doesn't just answer from memory but actually looks up the right documents first. In an enterprise setting, that knowledge base might contain internal policies, product documentation, regulatory circulars, customer support transcripts, or technical manuals. The retrieval pipeline finds the most relevant chunks, and the language model synthesizes an answer grounded in those sources.
The appeal for regulated industries is obvious. A bank doesn't want an AI system inventing loan eligibility criteria. A telecommunications provider doesn't want a chatbot making up data plan terms. RAG grounds responses in verifiable sources.
But here's where it gets complicated for Malaysian organizations: most RAG pipelines are built assuming English-language content and English-language queries. When your knowledge base contains a mixture of Bahasa Melayu, English, Manglish, and internal shorthand, the assumptions break down.
Why Generic RAG Pipelines Struggle With Malaysian Content
Universiti Teknologi MARA research on intent classification for Malay language queries highlights several persistent challenges: limited annotated datasets, severe class imbalance, and rich morphological variation. The study notes that "intent classification for Malay language queries remains a challenge due to limited annotated datasets, severe class imbalance, and rich morphological variation". This isn't a minor technical footnote. It means that a retrieval model fine-tuned on English Wikipedia will not automatically perform well when someone searches "cara nak mohon kad kredit" or "baki akaun saya tak update."
Bahasa Melayu has its own morphological rules, prefixes, suffixes, and compound structures that don't map cleanly onto English tokenization. A query like "pengeluaran wang tanpa kad" (cardless withdrawal) contains morphological derivations that a model unfamiliar with Malay might fragment incorrectly.
Manglish and Code-Switching Break Retrieval
Manglish isn't "bad English." It's a colloquial Malaysian variety that involves systematic code-switching between Bahasa Melayu, English, Chinese dialects, Tamil, and local slang. It has its own grammar, rhythm, and pragmatic conventions. IEEE research on Malay-English code-mixed text describes Manglish as a "Malay-English code-switched language commonly used in Malaysia and Singapore, where speakers seamlessly mix Malay and English within the same conversation".
The technical challenge is that code-switching creates retrieval problems. If a user asks "Boleh explain kenapa loan saya kena reject?" the query contains English (explain, loan, reject) and Malay (boleh, kenapa, saya, kena) interwoven. A monolingual embedding model may retrieve documents about "loan rejection" but miss Malay-language policy documents that discuss "penolakan pinjaman." Conversely, a Malay-focused model might miss English technical documentation.
Research on Manglish NLP notes that "the irregular syntax, informal spelling, and constant code-switching of Manglish hinder the performance of traditional Natural Language Processing systems trained on monolingual data".
Local Terminology and Enterprise Vocabulary
Every industry has its own vocabulary. In Malaysian banking, terms like "BPA" (Bank Perusahaan Kecil dan Sederhana), "CASA" (Current Account Savings Account), and "RENTAS" (Real-time Electronic Transfer of Funds and Securities) are standard. In telecommunications, abbreviations like "P1," "Unifi," and "postpaid plans" carry specific meanings. Healthcare has its own multilingual clinical terminology.
A generic RAG system has no way of knowing that "BPA" refers to SME banking and not a chemical compound. Enterprise vocabulary must be embedded into the retrieval pipeline through careful document preprocessing, terminology injection, and domain-specific evaluation.
Multilingual Embeddings: The Foundation of Retrieval Quality
Embeddings are mathematical representations of text that capture semantic meaning. When a user asks a question, the system converts that question into an embedding and searches for documents with similar embeddings. The quality of this retrieval step determines whether the AI generates a useful answer or a confident hallucination.
Research from Mesolitica, a Malaysian AI research group, demonstrates why generic embedding models fall short for Bahasa Melayu. The study notes that "when it comes to the Malay language, the performance of such out-of-the-box solutions falls short. The intricacies of Malay linguistics, along with its unique semantic structure, pose challenges that generic models struggle to overcome".
The researchers fine-tuned Llama2 models specifically for Malaysian semantic representations and found that their 600-million-parameter model outperformed OpenAI's text-embedding-ada-002 across recall metrics for Malaysian news, Twitter, and forum datasets. Their 2-billion-parameter model achieved superior Recall@5 and Recall@10 for the "Melayu" keyword research papers dataset.
This doesn't mean every enterprise needs to train its own embedding model. But it does mean that selecting an embedding model requires testing against your actual data -- your documents, your user queries, your language mix. A model that performs well on English benchmarks may underperform on Malay queries.
Multilingual vs. Language-Specific Approaches
There are several approaches enterprises can consider:
Multilingual embeddings: Models like multilingual-e5-large-instruct and BGE-M3 are trained across many languages and can handle queries in different languages without translation. Research shows that multilingual-e5-large-instruct achieves high similarity scores for Indonesian-Malay positive samples (0.9682). These models are practical when your knowledge base and user queries span multiple languages.
Language-specific embeddings: Models fine-tuned specifically for Bahasa Melayu or Malaysian content may outperform multilingual models on Malay-centric tasks. The trade-off is that they may perform poorly on English queries or code-switched content.
Fine-tuned models: Organizations with sufficient data can fine-tune embedding models on their own document corpus and query patterns. This is the most resource-intensive approach but can yield the best retrieval quality for organization-specific vocabulary.
There is no universally superior approach. The right choice depends on your data distribution, query patterns, privacy requirements, and infrastructure constraints.
Document Chunking and Preprocessing for Multilingual Content
How you prepare documents before embedding them matters as much as which model you choose. A 50-page regulatory circular contains sections in Bahasa Melayu with English technical appendices. A customer support transcript mixes Malay, English, and Manglish. A product manual has bullet points in English but safety warnings in Malay.
Naive chunking -- splitting documents by fixed token counts -- can separate a Malay paragraph from its English explanation, destroying semantic coherence. Better approaches include:
Language-aware chunking: Detect language boundaries and chunk accordingly, keeping semantically related content together regardless of language.
Semantic chunking: Use embedding similarity to identify natural break points rather than fixed token counts.
Metadata enrichment: Tag chunks with language, source, department, and topic metadata to enable filtered retrieval.
Terminology normalization: Create synonym mappings for enterprise abbreviations and local terminology so that "BPA" and "SME banking" retrieve the same documents.
Retrieval Quality and Reranking
Initial retrieval using vector similarity returns a set of candidate documents. Reranking improves precision by re-scoring those candidates using a more sophisticated model that considers the full query-document relationship.
For multilingual RAG, reranking becomes especially important. A query in Manglish might retrieve documents in both Malay and English. A reranker can evaluate which documents actually answer the question, regardless of language, and surface the most relevant results.
The RAG triad -- Context Relevance (is the retrieved context relevant?), Groundedness (is the answer supported by the context?), and Answer Relevance (does the answer address the question?) -- provides a framework for evaluation. For multilingual systems, language consistency becomes an additional metric: does the system respond in the language the user expects, or does it awkwardly mix languages?
Evaluation: The Hardest Part of Localized RAG
Evaluating RAG systems is difficult in any language. Evaluating them for Bahasa Melayu and Manglish adds layers of complexity.
Automated metrics like Recall@K, Precision@K, and Mean Reciprocal Rank measure retrieval quality but don't capture whether responses sound natural to Malaysian users.
Human evaluation becomes essential. Native speakers of Bahasa Melayu and Manglish need to assess whether responses are grammatically correct, culturally appropriate, and actually useful. A response that is factually accurate but written in formal textbook Malay may feel alienating to a user who asked in casual Manglish.
Language consistency is a unique challenge for multilingual RAG. Research on the MARS evaluation framework introduces Language Consistency as "a newly introduced metric to measure a unique challenge in multilingual RAG". The system should respond in the language of the query, or at least in a language the user is comfortable with, without jarring switches.
Domain-specific evaluation ensures that terminology is used correctly. A healthcare RAG system should know the difference between "demam denggi" (dengue fever) and "demam biasa" (common fever) and use the correct clinical terminology in responses.
Access Control and Data Security
In regulated industries, RAG systems handle sensitive information. A banking RAG system might retrieve documents containing customer data, internal policies, or regulatory correspondence. Access control must ensure that users only retrieve documents they are authorized to see.
This requires integration between the RAG pipeline and enterprise identity systems. Role-based access control filters must be applied at the retrieval stage, not just at the user interface. If a customer service representative doesn't have access to legal documents, the RAG system should never retrieve those documents, even if they contain relevant information.
Malaysia's Personal Data Protection Act amendments, effective through 2025, introduced a 72-hour breach notification requirement and direct liability for data processors. For organizations building RAG systems that process personal data, this means the pipeline must be designed with data minimization, encryption, and audit trails from the start.
A Practical Enterprise Scenario
Consider a telecommunications company that wants to build an internal knowledge assistant for its customer service team. The knowledge base contains:
- Product documentation in English and Bahasa Melayu
- Internal troubleshooting guides with technical English terminology
- Customer support transcripts in Manglish
- Regulatory circulars from the Malaysian Communications and Multimedia Commission
A customer service representative receives a query: "Customer complain line always drop, dah try restart router tapi still same. Apa nak buat?"
A generic RAG system might retrieve English troubleshooting guides that don't address the specific phrasing. A localized RAG system would:
- Preprocess the query, recognizing the code-switching between English (line, drop, restart, router) and Malay (complain, dah try, tapi, still same, apa nak buat)
- Retrieve documents in both languages, including Malay troubleshooting guides and English technical documentation
- Rerank results to prioritize documents that address "line drop" and "router restart" regardless of language
- Generate a response that acknowledges the customer's language register -- perhaps responding in a mix of Malay and English that feels natural
The system would be evaluated not just on whether it retrieved the right documents, but on whether the response actually helps the customer service representative solve the problem.
Building Localized RAG Systems: A Practical Path
Organizations exploring localized RAG for Bahasa Melayu and Manglish should consider these practical steps:
Start with data assessment. Survey your knowledge base. What languages are present? What's the ratio of formal to informal content? What enterprise terminology needs special handling?
Select embedding models empirically. Test multiple embedding models against your actual documents and queries. Measure retrieval quality for Malay, English, and code-switched queries separately.
Invest in preprocessing. Document chunking, language detection, and terminology normalization are not optional extras. They determine whether retrieval works at all.
Build human evaluation into the process. Native speakers must review system outputs for language quality, cultural appropriateness, and usefulness.
Design for access control from day one. Integrate identity and authorization systems into the retrieval pipeline.
Plan for continuous improvement. Language evolves. New products introduce new terminology. User query patterns change. The RAG system needs monitoring, feedback loops, and periodic re-evaluation.
Where Technical Expertise Meets Regulatory Reality
Malaysia's AI governance landscape is evolving. The proposed AI Governance Bill, expected to be presented to Cabinet in June 2026, proposes five guiding principles including transparency, accountability, and safe and secure use of AI systems. The Personal Data Protection Act amendments create direct liability for data processors, meaning AI vendors share regulatory exposure.
For Malaysian enterprises, building localized RAG systems isn't just a technical challenge. It's a governance challenge. The system must be designed, tested, documented, and maintained in ways that support regulatory requirements while remaining useful to actual Malaysian users.
This is where specialized expertise matters. Evalogical's AI and LLM solutions support organizations in designing, integrating, testing, and maintaining enterprise AI systems. From custom software development to AI testing and evaluation, their capabilities address the full lifecycle of localized RAG deployment -- not just the initial build, but the ongoing evaluation, monitoring, and improvement that keeps systems useful as language and enterprise needs evolve.
Frequently Asked Questions
What is localized RAG and how does it differ from standard RAG?
Localized RAG adapts the standard Retrieval-Augmented Generation pipeline to handle language variations, code-switching, local terminology, and culturally specific content. While standard RAG assumes relatively uniform language input, localized RAG addresses the reality that Malaysian enterprise content and user queries often mix Bahasa Melayu, English, Manglish, and industry-specific vocabulary. It's not simply translating an English RAG system into Malay -- it involves different embedding models, preprocessing strategies, evaluation approaches, and response generation.
Can I just translate my English documents into Bahasa Melayu and use a standard RAG system?
Translation-based approaches have significant limitations. Machine translation often misses cultural context, local terminology, and the nuances of code-switching. More importantly, Malaysian users don't always query in formal Bahasa Melayu -- they use Manglish, abbreviations, and mixed-language expressions. A translation-based RAG system would need to handle these variations in the query itself, not just the documents. Retrieval using multilingual embeddings, where the system searches across language boundaries without explicit translation, is often more practical.
What embedding models work best for Bahasa Melayu and Manglish?
There is no single best model. Research from Mesolitica shows that fine-tuned Malaysian-specific models can outperform general multilingual models on Malay retrieval tasks. Multilingual models like multilingual-e5-large-instruct and BGE-M3 perform well across languages including Malay. The right choice depends on your data distribution, query patterns, and whether you need to handle English and Malay queries equally or primarily one language. Empirical testing against your actual corpus is essential.
Does RAG eliminate hallucinations in enterprise AI systems?
No. RAG reduces hallucinations by grounding responses in retrieved documents, but it does not eliminate them. The system can still generate incorrect answers if retrieval fails, if the retrieved documents are ambiguous, or if the language model misinterprets the context. This is why evaluation -- both automated and human -- is critical. For Bahasa Melayu and Manglish systems, human evaluation by native speakers is especially important because automated metrics may not capture language quality or cultural appropriateness.
How do I evaluate a multilingual RAG system for Malaysian languages?
Evaluation should combine automated retrieval metrics (Recall@K, Precision@K), generation metrics (groundedness, answer relevance), and language consistency measures. Human evaluation by native Bahasa Melayu and Manglish speakers is essential for assessing grammatical correctness, cultural appropriateness, and usefulness. Domain-specific evaluation ensures that enterprise terminology is used correctly. The RAG triad -- Context Relevance, Groundedness, and Answer Relevance -- provides a framework, with Language Consistency added as a metric specific to multilingual systems.
What are the data privacy considerations for RAG systems in Malaysia?
Malaysia's Personal Data Protection Act amendments, effective through 2025, introduced a 72-hour breach notification requirement and direct liability for data processors. If your RAG system processes personal data, you need to design for data minimization, encryption, access control, and audit trails. Retrieval should respect role-based access controls, ensuring that users only retrieve documents they're authorized to see. If using cloud-based AI services, data residency and cross-border transfer restrictions under the PDPA must be evaluated.
Ready to Build AI That Actually Understands Malaysian Language?
Building a RAG system that works for Bahasa Melayu and Manglish is a technical, linguistic, and governance challenge. It requires the right embedding models, careful preprocessing, rigorous evaluation, and a deep understanding of how Malaysian users actually communicate.
Evalogical's custom software development and AI integration services support Malaysian enterprises through every stage of this journey -- from initial data assessment and pipeline design to testing, evaluation, and ongoing monitoring. If your organization is exploring localized RAG for customer support, internal knowledge management, or regulatory compliance, the right technical foundation makes all the difference.
Your Trusted Software Development Company