I got tired of retrieval papers claiming their pipeline is "state-of-the-art" without significance tests, without telling you what dataset they ran on, and without reproducible code. So I built raven-retrieval -- 19 retrieval pipelines, real BEIR datasets, Bonferroni-corrected bootstrap signific...
Source: [Dev.to](https://dev.to/subhansh/i-benchmarked-19-retrieval-pipelines-head-to-head-and-the-results-were-surprisingly-honest-3efe)