Anthropic’s Opus 5 Shows Strong Resistance to Prompt Injection
Boris Cherny highlights Opus 5 as Anthropic’s least prompt-injectable model yet, based on internal evaluations and red teaming.
author
Baris operates SummarizedData and maintains its sources, publishing rules, and automation. Read the editorial process.
Boris Cherny highlights Opus 5 as Anthropic’s least prompt-injectable model yet, based on internal evaluations and red teaming.
Anthropic released Claude Opus 5, a proactive model that approaches frontier intelligence while costing half as much as Claude Fable 5 and topping the Artificial Analysis leaderboard.
a fallen power line caused 3 gigawatts of data centers to disconnect almost simultaneously, revealing risks to grid stability as ai demand grows.
openai's first hardware, the micro keypad, pairs with chatgpt and codex, offering customizable keys and voice dictation for $230, but faces mixed reviews.
prentis, an ai lab building computer-use agents for office tasks, is in talks to raise $100 million at a $1 billion valuation.
prentis, co-founded by reid hoffman and mark pincus, is in talks to raise $100 million at a $1 billion valuation to build ai agents that automate routine office workflows.
cognition gives devin a personality, anthropic launches opus 5, and ai security concerns grow after kimi k3 and openai breach.
cognition acquires poke to make its coding assistant devin more personable and proactive, reflecting a shift where ai interaction style becomes a competitive edge.
tech companies urge u.s. policymakers to avoid sweeping restrictions on open-weight ai models amid china ip theft concerns.
chinese model kimi k3 sparked u.s. industry alarm while an openai model breach at hugging face exposed broader security risks.
anthropic releases opus 5, a cheaper and less restrictive model that outperforms fable 5 on several benchmarks.
a transformer-based diffusion model improves imputation and forecasting of sparse hydrological data across multiple sites.
bluesky's ai assistant attie now lets users ask open-ended questions to research trends and influential accounts across the at protocol network.
a breakdown of the five engineering ideas that make agentic ai systems work in production, from tool use to evaluation.
a practical look at grapheval, a framework using knowledge graphs and nli to pinpoint hallucinations in language model outputs.
repeated sampling from one model at high temperature does not produce the structured diversity seen across different models, limiting its use for epistemic uncertainty.
a new retrieval system combines text, keyword, knowledge graph, and image signals to improve question answering over complex pdf collections.
a new benchmark measures how well large language models construct and evaluate training data by fine-tuning base models on their outputs.