dpo cuts text degeneration in ocr models
direct preference optimization reduced text degeneration by an average of 59.4% across five ocr model families by using the model's own failure outputs as rejection pairs.
author
Baris operates SummarizedData and maintains its sources, publishing rules, and automation. Read the editorial process.
direct preference optimization reduced text degeneration by an average of 59.4% across five ocr model families by using the model's own failure outputs as rejection pairs.
a new framework lets human agents approve or reject algorithmic price suggestions, using old pricing data to skip the slow start typical in sparse booking markets.
uber limits employee spending on agentic coding tools like claude code and cursor to $1,500 monthly per tool after overshooting its 2026 ai budget.
a new method reduces the cubic complexity of gaussian processes with gradients by using exact gradient reduction and vecchia approximation.
a position paper argues that in high-dimensional settings, many different mechanisms can produce the same data, so predictive success does not prove a model has found the true mechanism, and large language models can hide this by giving a single fluent explanation.
aura-mem uses a learned gate to write only when observations change actions, keeping memory fixed at 4,224 bytes regardless of episode length.
periodic and soft target updates can guarantee convergence in linear q-learning under explicit spectral and step-size conditions.
a new method uses gradient tests instead of validation loss to decide when to stop training gradient boosted trees, avoiding the need for a patience parameter.
a new diagnostic reveals that common anomaly detection benchmarks become unreliable when held-out classes overlap with normal data in representation space.
a new framework aligns structured electronic health record representations with large language models to improve clinical prediction and reasoning.
data security startup cyera is finalizing a $300 million round at a $12 billion valuation, even as it spends more than it earns.
microsoft announced two new language models, mai-thinking-1 and mai-code-1-flash, with low active parameter counts and claims of clean training data.
microsoft introduces scout, a persistent ai assistant built on the openclaw framework, offering customizable skills and security features for microsoft 365 users.
holo3.1 expands computer-use ai to mobile, desktop, and web with quantized models for local inference.
an overview of advances in making large language models more interpretable through dynamic evaluation, statistical methods, and accessible tools.
linkedin speeds up optimization with pytorch gpus, microsoft tests ai behavior from text, and google adds fake call detection to android.
google's new android feature silently verifies calls to stop ai voice impersonation scams.
microsoft's open source assert framework uses ai to turn plain-language rules into scored tests for application-specific ai behavior.