Entity

Self-Trained Verification for Training- and Test-Time Self-Improvement

Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall when verifier scores inflate while accuracy stagnates, and when feedback is too generic to act on; self-training fails similarly when bad self-generated data are added to training. Better verificatio

Paper · arXiv

cs.LG

Authors: Chen Henry Wu, Aditi Raghunathan
Published: 2026-05-28
Categories: cs.LGcs.AIcs.CL

Abstract ↗

via arXiv · 2605.3029