SISAP 2026 Indexing Challenge
Overview
The SISAP 2026 Indexing Challenge benchmarks approximate nearest-neighbor and sparse retrieval algorithms at scale. Participants submit a GitHub repository containing a Docker-based solution; the evaluation pipeline runs each solution on the official datasets and records build time, query time, and recall.
Three tasks are evaluated:
| Task | Description | Dataset(s) | Required recall |
|---|---|---|---|
| Task 1 | All-k-NN on dense embeddings (k = 15) | wikipedia-small, wikipedia-dev (public), wikipedia-eval (private test set; now published) | ≥ 0.80 |
| Task 2 | k-NN query search on dense embeddings (k = 30) | llama-dev (public), llama-eval (private test set; now public), llama-pg174 (test variant; now public) | ≥ 0.80 |
| Task 3 | k-NN query search on sparse embeddings (k = 30) | fiqa-dev (public), nq-eval (private test set; now public) | ≥ 0.90 |
The throughput metric is:
- Task 1: total vectors ÷ (build time + query time)
- Tasks 2 & 3: number of queries ÷ query time
Query counts are 6,000,000 (wikipedia-eval), 6,350,000 (wikipedia-dev), 200,000 (wikipedia-small), 1,000 (all three llama sets), 3,452 (nq-eval) and 6,648 (fiqa-dev).
Two leaderboard views
The Leaderboard and Plots pages offer a Scope switch:
- SISAP archived — the citable record accompanying the challenge overview paper, restricted to teams that contributed a system-description paper, plus the organizers’ baselines. These are the official rankings reported of the SISAP 2026 Indexing Challenge, and the default view until the conference in October 2026.
- Public leaderboard — every submitted run, including teams that did not participate in the SISAP 2026 Indexing Challenge. Stays open to new submissions.
For Task 1, wikipedia-small and wikipedia-dev are the public datasets, and wikipedia-eval was the private test set during evaluation.
For Task 2, llama-dev is the public dataset, llama-eval was the private test set during evaluation, and llama-pg174 is a variant of the test set.
For Task 3, fiqa-dev is the public dataset and nq-eval was the private test set during evaluation.
See the Leaderboard for current results and the Plots page for recall/QPS Pareto frontiers.
Datasets
All datasets are now public in the SISAP 2026 dataset collection on Hugging Face.
The private test workloads were created in the following way:
- Task 1: See https://github.com/sisap-challenges/sisap26-leaderboard/tree/main/workload_generation.
- Task 2: We replaced the vectors with a random subset of unseen vectors from https://huggingface.co/datasets/vector-index-bench/vibe/blob/main/README.md.
- Task 3: OpenAI’s ChatGPT Pro 5.4 was used to paraphrase all queries in the corpus before embedding. Prompt: “Paraphrase all questions in nq/queries.jsonl”, after which the modified queries were run through the same embedding model.
Credits
This leaderboard used the evaluation website of VIBE: Vector Index Benchmark for Embeddings by Jääsaari et al. as a starting point.