Author: Bruyn, Maxime De; Lotfi, Ehsan; Buhmann, Jeska; Daelemans, Walter
                    Title: MFAQ: a Multilingual FAQ Dataset  Cord-id: qa9ez4lc  Document date: 2021_9_27
                    ID: qa9ez4lc
                    
                    Snippet: In this paper, we present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages. Although this is significantly larger than existing FAQ retrieval datasets, it comes with its own challenges: duplication of content and uneven distribution of topics. We adopt a similar setup as Dense Passage Retrieval (DPR) and test various bi-encoders on this dataset. Our experiments reveal that a multilingual model based on XLM-RoBERTa ach
                    
                    
                    
                     
                    
                    
                    
                    
                        
                            
                                Document: In this paper, we present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages. Although this is significantly larger than existing FAQ retrieval datasets, it comes with its own challenges: duplication of content and uneven distribution of topics. We adopt a similar setup as Dense Passage Retrieval (DPR) and test various bi-encoders on this dataset. Our experiments reveal that a multilingual model based on XLM-RoBERTa achieves the best results, except for English. Lower resources languages seem to learn from one another as a multilingual model achieves a higher MRR than language-specific ones. Our qualitative analysis reveals the brittleness of the model on simple word changes. We publicly release our dataset, model and training script.
 
  Search related documents: 
                                Co phrase  search for related documents- locality sensitive hashing and lsh locality sensitive hashing: 1, 2
 
                                Co phrase  search for related documents, hyperlinks ordered by date