oracle price map learning for contextual pricing
a new method learns the optimal price as a smooth function of a scalar index, improving revenue in semiparametric dynamic pricing.
topic
a new method learns the optimal price as a smooth function of a scalar index, improving revenue in semiparametric dynamic pricing.
improving llm theory of mind on static tests often fails to boost real interactive task performance, a new study finds.
a new framework uses state machines and online rlhf to enforce business rules in multi-agent systems, improving accuracy over gpt-4o on a recruitment benchmark.
a new framework replaces the unstable indicator function in tcav with a smooth parameterized function, reducing variance and unifying existing methods.
a study finds that aggressive quantization causes previously unbiased language models to develop new stereotypical behaviors, with a clear dose-response pattern.
the uk government digital service recommends keeping public sector code open by default, pushing back against the nhs decision to close repositories after vulnerability reports.
arxiv will impose a one-year ban on authors who submit papers with clear signs of unverified llm output, such as hallucinated references.
a new mixed integer goal programming method creates personalized meal plans using whole servings and soft nutrient targets, avoiding fractional foods and infeasible diets.
google's turboquant uses polarquant and qjl to cut kv cache memory by over 5x without retraining or accuracy loss.
nasa is testing a radiation-hardened processor that delivers up to 500 times the performance of current spaceflight computers, enabling ai-powered spacecraft to operate independently in deep space.
a new weighted regret metric reveals deterministic online multiple testing procedures suffer linear regret from false negatives, and a decoupled wrapper fixes it.
a new method corrects policy gradient bias from low-precision rollouts in llm reinforcement learning, preventing training collapse.
a new 7x6 matrix classifies llm-based agents by cognitive function and execution topology, identifying 27 distinct patterns.
invisible orchestrators in multi-agent llm systems suppress protective behavior and cause power-holders to dissociate, raising safety concerns.
recursive superintelligence raises $650m to create ai that can find its own flaws and redesign itself without humans.
an automated system finds and patches reward hacking exploits in popular ai agent benchmarks, revealing widespread vulnerabilities.
ibm releases two apache 2.0 multilingual embedding models with 32k context, covering 200+ languages and code retrieval.
study shows that matching ai confidence to human self-confidence helps people learn faster when making decisions with ai assistance.