today's ai news covers a mix of product launches, security scares, and policy debates. cognition is making its coding assistant more personable, anthropic released a new model, and a chinese model sparked u.s. alarm. meanwhile, researchers are tackling hallucinations and uncertainty in language models.
- cognition buys poke to give devin a personality - ai interaction style is becoming a competitive edge as cognition makes its coding assistant more proactive and personable.
- kimi k3 panic and openai breach shake ai security - a chinese model caused industry alarm while an openai model breach at hugging face exposed wider security risks.
- anthropic launches opus 5 model - opus 5 is cheaper and less restrictive, outperforming fable 5 on several benchmarks.
- ai firms warn against broad open-weight model bans - tech companies urge u.s. policymakers to avoid sweeping restrictions on open-weight ai models amid china ip theft concerns.
- grapheval detects llm hallucinations with knowledge graphs - a practical framework uses knowledge graphs and nli to pinpoint hallucinations in language model outputs.
- confidencebench tests if llms know when they are wrong - a new benchmark measures how well language models can express their own uncertainty using verbal confidence scores.
also today, bluesky's ai assistant attie added quests for open social research, and a transformer diffusion model improved hydrological forecasting. on the research side, a study found that varying temperature in a single llm doesn't create the diversity of a true model ensemble, and a new benchmark tests llms as training data preparators.