Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
International Conference on Learning Representations, 2026
Core idea. Uses adversarial multi-agent debate to fill missing relevance judgments in IR benchmarks, auto-labeling agreement cases and escalating only disagreements to human annotators.