source: Hugging Face Blog: Your Agent Aced the Task. Will It Do It Again?
level: technical
a react agent using gpt-4.1 scored 77.4% average success on appworld tasks across five runs. but it succeeded on all five runs for only 53.0% of tasks. that 24.4-point gap means many tasks work sometimes and fail other times with nothing changed. the authors built a consistency analyzer to find decision steps where the model was one token-sample away from doing something different. it needs one recorded trace and no ground truth, resampling each decision point with five completions.
the analyzer flags flat probability distributions where near-tied tokens can flip under small platform noise. from those flagged steps, it generates consistency guidelines in the altk-evolve format. on appworld test_normal with 168 tasks, adding these guidelines raised pass^5 from 53.0% to 69.0% while mean@5 rose from 77.4% to 81.0%. the consistency gap shrank from 24.4 to 12.0 points. medium and hard tasks gained about 45% relative improvement. mean accuracy never dropped.
the guidelines also transferred to similar tasks, lifting pass^5 by 13.0 points on related variants. a weaker model, gpt-oss-120b, saw pass^5 rise from 10.1% to 16.1% on same tasks and 8.7 points on similar tasks. the method uses one extra llm call per decision step, making it usable on production traffic where full replays are impossible. the code is open source in the altk-evolve repository, with a technical report on arxiv.
why it matters: for ai agents in finance or contracts, a task that works once but fails on repeat runs is a reliability problem that average accuracy hides; this method measures and reduces that gap.
source: Hugging Face Blog: Your Agent Aced the Task. Will It Do It Again?