AI & Machine Learning United States AI Technology Infrastructure Analysis VPO Test-time Search Reward Vectors GRPO Evaluators
May 2026

Reward Vectors Feed Search

VPO turns verifier-legible diversity into a post-training target, shifting leverage from single-answer alignment to reward-vector design.

If VPO is right, the next AI training advantage is not a bigger scalar reward model but ownership of decomposed reward vectors that keep useful alternatives alive for search.

A May 2026 arXiv preprint, Vector Policy Optimization, attacks a deployment mismatch: LLMs are increasingly used as generators inside best-of-k and evolutionary search loops, while post-training still often optimizes a fixed scalar reward that narrows the response distribution. VPO replaces scalar convergence with a set-level objective: generate multiple candidates, sample reward weightings over a vector-valued reward, and reward coverage of the Pareto front rather than collapse to one mode [1].

The asymmetric signal is not “diversity is good.” It is that verifier-legible diversity becomes inventory: a model trained to preserve different reward tradeoffs may be more valuable inside search systems than one trained to maximize a single average score [1] [2].

0.832
MuSiQue Best@30
VPO’s best@30 on the 300-question MuSiQue split, versus 0.728 for scalar GRPO.
Compute Gap
GRPO/GDPO with three times the rollouts and LM compute still stayed below VPO at n=8 on MuSiQue.
32
Hard LCB Problems
In the LiveCodeBench/OpenEvolve case study, VPO kept finding solutions over 200 iterations while GRPO plateaued early.
UNLOCK ACCESS

Want to read the full Signal?

Sign up to access the complete analysis for Reward Vectors Feed Search.