what reward learning gets from best-of-n preference data
analysis reveals how bradley-terry reward models trained on best-of-n preference data converge to specific targets depending on n and the base distribution.
topic
analysis reveals how bradley-terry reward models trained on best-of-n preference data converge to specific targets depending on n and the base distribution.
a multi-model study shows that fine-tuned deceptive language models develop early, linearly detectable representations of dishonesty.
a project runs python asgi web apps entirely in the browser using pyodide and a service worker, removing the need for a backend server except for static files.
new theory reveals when momentum helps or hurts in sparse training settings, based on two key timescales.
a metr study found most developers won't work without ai, but research suggests ai may slow them down and increase maintenance costs.
a plain-language glossary of common ai terms like llm, rag, and hallucination for anyone who nodded along but wants clarity.
a new method uses deep neural networks and adaptive prediction-powered learning to optimize treatment rules for bivariate survival outcomes in randomized trials.
a study of acl rolling review papers finds llm reviews align only moderately with human reviews and can be gamed by authors using iterative revision.
a 306m-parameter transformer with simplicial message passing reduces perplexity by 12% over gpt-2 small on wikitext-103.
a new framework uses labeled data from related tasks to improve statistical power in prediction-powered inference when only a handful of labels are available per task.
a compact binary mask reverses most knowledge edits in language models, revealing a shared mechanism behind diverse factual updates.
a new mirror-prox temporal-difference method uses behavior-policy transition information instead of feature covariance to speed up off-policy prediction.
new lower bounds show that the bandwidth term in federated probe-logit distillation is tight, and the method extends to nodes with different upload budgets.
a new llm-based agent called trace uses tool planning to optimize drug-like molecules over multiple steps, improving properties while keeping key structures intact.
replacing the auxiliary covariance matrix with the behavior bellman matrix improves stability in off-policy temporal-difference learning.
a new method extends federated conformal rag to provide valid uncertainty estimates at any stopping time, enabling safer adaptive control in distributed language model systems.
itbench-aa evaluates ai agents on kubernetes incident response, with claude opus 4.7 leading at 47% accuracy.
google research highlights from i/o 2026 include new ai tools for scientific discovery, health coaching, and edge computing, plus advances in weather prediction and model factuality.