olmoearth v1.1 cuts compute costs up to 3x
a new family of remote sensing models reduces token sequence length to lower inference costs while matching previous performance.
topic
a new family of remote sensing models reduces token sequence length to lower inference costs while matching previous performance.
gemini omni flash can generate and edit videos from text, images, audio, and video inputs, starting with the gemini app and youtube shorts.
google introduces experimental ai tools to speed up hypothesis generation, computational discovery, and literature review for researchers.
google research's empirical research assistance (era) uses gemini to automate and optimize scientific coding, now published in nature and powering the computational discovery prototype.
penn researchers created exciton-polaritons that switch signals with tiny energy, potentially enabling faster, more efficient photonic ai computing.
a new method uses counterfactual trajectory comparisons to turn sparse terminal rewards into step-level signals, stabilizing reinforcement learning for multi-step llm reasoning.
attackers who delete legal actions before an agent decides cause severe and lasting damage across multiple games and algorithms.
signmuon combines 1-bit sign communication with muon's matrix-aware updates to reduce distributed training bottlenecks.
skim uses site-specific templates to skip heavy ai steps for most web tasks, falling back to full agents only when needed.
agentwall intercepts agent actions before execution, enforcing declarative safety policies to prevent harmful operations on local machines.
agentstop reduces energy use in local ai agents by predicting task failure early, cutting token waste and battery drain on consumer devices.
a new method learns safe decision rules from logged data using general risk measures like cvar, with strong theoretical guarantees.
a lightning talk summary of major llm developments from november 2025 to may 2026, including coding agent breakthroughs and open-weight model advances.
sandboxaq integrates its physics-grounded quantitative models into anthropic's claude, letting researchers run complex simulations through plain language prompts.
instruction-tuned language models show fair outputs but retain biased internal representations that can reverse decisions when activated.
a trust-region method for fine-tuning multi-agent llm teams avoids compounding errors from stale rollouts, outperforming baselines by 7.1%.
a new open benchmark evaluates complete agent systems, not just models, across six diverse tasks to measure generality, quality, and cost.
deepslide is a multi-agent system that helps prepare entire presentations, from planning and slide creation to rehearsal support.