Papers
arxiv:2506.01615

IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems

Published on Jun 2, 2025
Authors:
,
,
,

Abstract

The development of RAG systems for Indian languages is supported by creating a multilingual benchmark (IndicMSMarco) and a large-scale training dataset of (question, answer, relevant passage) tuples from Wikipedias in 19 Indian languages.

Retrieval-Augmented Generation (RAG) systems enable language models to access relevant information and generate accurate, well-grounded, and contextually informed responses. However, for Indian languages, the development of high-quality RAG systems is hindered by the lack of two critical resources: (1) evaluation benchmarks for retrieval and generation tasks, and (2) large-scale training datasets for multilingual retrieval. Most existing benchmarks and datasets are centered around English or high-resource languages, making it difficult to extend RAG capabilities to the diverse linguistic landscape of India. To address the lack of evaluation benchmarks, we create IndicMSMarco, a multilingual benchmark for evaluating retrieval quality and response generation in 13 Indian languages, created via manual translation of 1000 diverse queries from MS MARCO-dev set. To address the need for training data, we build a large-scale dataset of (question, answer, relevant passage) tuples derived from the Wikipedias of 19 Indian languages using state-of-the-art LLMs. Additionally, we include translated versions of the original MS MARCO dataset to further enrich the training data and ensure alignment with real-world information-seeking tasks. Resources are available here: https://huggingface.co/datasets/ai4bharat/Indic-Rag-Suite

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2506.01615
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2506.01615 in a model README.md to link it from this page.

Datasets citing this paper 12

Browse 12 datasets citing this paper

Spaces citing this paper 29

Browse 29 spaces citing this paper

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.