
RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models
Meta researchers present RankEvolve, a multi-agent auto-research framework that evolves a generative ranking model end to end. An Executable Operating Protocol compiled into a runtime-enforced state machine cures long-horizon drift; a meta-meta-harness binds complete coding-agent products (Claude Code, Codex) as graph nodes that cross-check each other; and a knowledge layer carries findings and negative results across iterations. On the open-source HSTU recommender it reported NDCG@10 0.2192 on MovieLens-20M LARGE (+4.48% over the published anchor). A hidden-oracle benchmark (ExecML) shows the CC+Codex pair lifts all-oracle execution accuracy from 45.8% to 62.5% at matched budget (+16.7, 95% CI [6.6,26.7]) while cutting the silent critical-defect rate to 10.4%, and the effect replicates on LitGPT (+12.5).