why few-step text latents fail but image latents work
few-step text generation fails because of sharp categorical readouts, not poor transport, as proven by a geometric analysis of continuous text decoders.
topic
few-step text generation fails because of sharp categorical readouts, not poor transport, as proven by a geometric analysis of continuous text decoders.
scarfbench evaluates ai agents on real-world enterprise java framework migrations, measuring build, deploy, and behavioral success.
a new framework edits internal representations during reasoning to guide large language models toward truthful answers without disrupting correct thinking.
every eval ever and hugging face community evals are now intercompatible, enabling cross-posting and interpreting evaluation results with links to open models and leaderboards.
google research releases building-level rooftop reflectivity data for over 50 global cities to help planners target cool roof interventions and reduce urban heat.
miles is an open source framework that combines sglang, megatron-lm, ray, and pytorch for scalable reinforcement learning post-training of large language models.
google deepmind launches nano banana 2 lite for fast image generation and gemini omni flash for video generation and conversational editing.
a new online algorithm for high-dimensional quantile regression uses delayed hard thresholding to find sparse models from streaming data.
optimization theory, biology, markets, and machine learning all point to the same conclusion: under finite resources, focused systems outperform general ones.
a study shows how anti-symmetric perturbations can reduce variance in stochastic gradient langevin monte carlo estimators under small step sizes.
a new benchmark evaluates how well ai models generate scientific figures with correct labels, relations, and conventions.
a new method initializes sigmoidal neural network weights using spectral geometry from data, improving training and performance.
a position paper argues that reinforcement learning researchers should clearly distinguish between optimizing for simulator performance and using simulators as stand-ins for real-world deployment.
gptnt tests how well ai agents communicate under time pressure using the game keep talking and nobody explodes.
deepreinforce releases ornith-1.0, an open weights model family for agentic coding built on gemma 4 and qwen 3.5, achieving top open-source coding benchmarks.
a new method links model failures to specific data fixes using capability slices, making pre-training optimization systematic instead of guesswork.
a study compares self-evolving llm agents and finds that recursive rewriting with a held-out gate improves robustness across benchmarks.
a new ai framework uses imaging data to measure galaxy distances with near-spectroscopic accuracy, aiding dark energy studies.