source: arXiv Artificial Intelligence: Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

level: research

Researchers introduced RubricForge, a method that induces a judging rubric from a small set of ground-truth-labeled trajectories. The rubric is evolved through reflective evolution to maximize agreement with environment rewards, then frozen and applied to held-out trajectories in a single model call without environment access. The optimized artifact is human-readable text, so every verdict can be inspected. This approach targets the problem of over-crediting fluent but unsuccessful agent trajectories, which plagues existing judges that hand-write rubrics or fine-tune judge weights.

The paper reports that RubricForge reduces over-crediting compared to baseline judges. In experiments, it improved agreement with true environment rewards on held-out trajectories. The method requires only a small set of labeled examples, making it practical when gold signals are scarce. The induced rubric is frozen after evolution, so it does not adapt during evaluation. This design choice trades flexibility for stability and interpretability. The authors note that the rubric's quality depends on the diversity and correctness of the labeled trajectories used for induction.

Automatic judges are increasingly used to evaluate language-model agents when executable environment rewards are unavailable or too costly. Existing approaches like G-Eval rely on hand-written rubrics, which may not capture all failure modes. Fine-tuning judge weights can also overfit to superficial fluency. RubricForge offers a middle path: it learns a rubric from data but keeps it as text, enabling human review. This could improve trust in agent evaluations and reduce the need for expensive environment interactions during deployment.

why it matters: More reliable automatic judges can reduce the cost and time of evaluating AI agents while avoiding false confidence in unsuccessful systems.


source: arXiv Artificial Intelligence: Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation