adaptive importance sampling fixes quantized rl training
a new method corrects policy gradient bias from low-precision rollouts in llm reinforcement learning, preventing training collapse.
aisummaries filed under research
a new method corrects policy gradient bias from low-precision rollouts in llm reinforcement learning, preventing training collapse.
aia new 7x6 matrix classifies llm-based agents by cognitive function and execution topology, identifying 27 distinct patterns.
aiinvisible orchestrators in multi-agent llm systems suppress protective behavior and cause power-holders to dissociate, raising safety concerns.
airecursive superintelligence raises $650m to create ai that can find its own flaws and redesign itself without humans.
aian automated system finds and patches reward hacking exploits in popular ai agent benchmarks, revealing widespread vulnerabilities.
aiibm releases two apache 2.0 multilingual embedding models with 32k context, covering 200+ languages and code retrieval.
aistudy shows that matching ai confidence to human self-confidence helps people learn faster when making decisions with ai assistance.
aia new model uses physics-based concepts to make ocean heat forecasts interpretable, revealing the drivers behind predictions.
aia new method learns both control and when to communicate, using a safety shield to reduce sampling while maintaining stability.
aia quantum-inspired algorithm simulates complex quasicrystals with over 268 million sites, enabling faster design of advanced quantum materials.
aia new theory shows that fine-tuning a strong model on a weak model's outputs can elicit pre-trained knowledge without losing general skills.
ainew methods produce nested prediction sets across multiple coverage levels for simultaneous uncertainty quantification in online settings.
aia new framework learns latent user preferences from few conversations and applies them across tasks to improve human-aligned decision making.
aia test-time framework uses a generative verifier to pick the best action from multiple candidates, improving robustness of multimodal language model agents without retraining.
aia new metric called quide collapses compression, accuracy, and latency into one score to find optimal quantization levels for neural networks.
aiai analysis of natural speech patterns like pauses and filler words can predict executive function and may help detect early cognitive decline.
aiuniform control schedules hurt text quality in discrete diffusion models, but a new adaptive scheduler improves multi-attribute steering by aligning interventions with each attribute's unique denoising timeline.
aia new post-hoc layer gives frozen predictors a spatial view of errors and a closed-form covariance without retraining.
aia new framework enables valid statistical inference when reusing data from adaptive sampling methods like bayesian optimization.
ailarge-depth transformers trained with adamw converge uniformly to a forward-backward ode system, with explicit convergence rates.
ai