This week saw major releases and research across AI, from Anthropic's watermarking plans for Claude to Google DeepMind's Gemini 3.7 Flash and sign language AI on Pixel phones. Meanwhile, studies questioned how well large language models handle research integrity and factual recall, and Hugging Face data revealed shifts in open model development toward China, small models, and coding agents.

  1. Anthropic details how Claude's new watermarks will work - Anthropic will use Google DeepMind's SynthID-Text to watermark Claude outputs, complying with the EU AI Act. The method survives editing and works for code, with detection available through a verifier.
  2. Google DeepMind Launches Gemini 3.7 Flash - Gemini 3.7 Flash is a more intelligent and cost-effective model for coding and agent workflows, continuing Google's rapid release cadence.
  3. Open models shift toward China, small models, and agents - Hugging Face data from early 2026 shows Chinese labs dominating frontier open models, small models driving downloads, and coding agents becoming a major user base.
  4. Benchmark Reveals LLMs Fail One in Three Research Integrity Decisions - IntegrityBench evaluated 18 frontier models on research integrity under pressure, finding significant failure rates and inconsistent behavior.
  5. Recall, not encoding, limits LLM factuality - Google Research's knowledge profiling shows frontier LLMs encode nearly all facts but struggle to recall them, with thinking recovering many failures.
  6. Google DeepMind brings sign language AI to Pixel phones - Google DeepMind's SL2T model powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with ASL to English.
  7. Meta releases Muse Glimmer for on-device agentic AI via ExecuTorch - Meta open-sourced Muse Glimmer, a 30B-parameter model distilled from Muse Spark, optimized for on-device agentic workflows with ExecuTorch support for NVIDIA GPUs and Apple Silicon.

The strongest shared signal is the push toward practical deployment and accountability: watermarking for compliance, on-device models for agents, and benchmarks that expose weaknesses in integrity and factuality. These efforts reflect a maturing field focused on real-world reliability.