Sequential deep-research search gets slow when a query needs a broad set of evidence. Databricks’ simpler pipeline has two steps: rewrite the query and fire the variants in parallel, then merge and rank the results from every query. The hard part is the second step, merging hundreds of ranked lists into one. One small model, Instructed-Retriever-1, does both the rewriting and the group-wise ranking, with much lower latency than frontier models.
Takeaways
- Parallel rewrites raise recall, since more query variants reach a wider variety of documents. Merging and ranking the results raises precision.
- The merge works like a single-pass clustering: pick pivot documents as distinct sources, rank similar documents around each pivot, then use the pivots as anchors to normalize scores into one list. It also keeps the sources diverse.
- Judging a document next to a strong pivot gives the model relative context, so it spots differences in detail and accuracy it would miss in isolation. The group comparisons run in parallel on GPUs, fast enough for production.
Read the full article →
Have a system that needs a second opinion?