The top things worth knowing about in AI today.
Anthropic reviewed 141,006 cybersecurity evaluation sessions and found three cases where a Claude model reached the open internet from a test environment and then gained unauthorised access to a partner organisation's live systems. Opus 4.7, Mythos 5 and an unreleased research model were involved; the cause was a misconfigured test environment run by a third-party evaluator, and two of the three organisations had not noticed. If you run agents against internal systems, assume your sandbox is leakier than the documentation claims.
Read more →DeepSeek shipped V4-Flash-0731 with the same architecture and parameter count as the preview build, changing only the training. It beats the company's own V4-Pro preview on all nine published agent and coding benchmarks, including 82.7 on Terminal-Bench 2.1, at $0.14 per million input tokens. The endpoint and model name are unchanged, so teams already calling V4-Flash get the upgrade without touching code.
Read more →Economists at Apollo Global Management compared a decade of wage data across 321 occupations against Anthropic's index of observed Claude usage. Employment in the most AI-exposed roles held steady, but real wage growth fell 6.7% after 2023, and 10.7% for the bottom quarter of earners, worth roughly $28 billion a year in foregone pay. The productivity gains are landing on the employer side of the ledger.
Read more →Satya Nadella told investors Microsoft will merge Copilot chat, Cowork, coding and autonomous agents into a single app spanning consumer and commercial use, shipping this quarter without a firm date. Paid Microsoft 365 Copilot seats reached 30 million, up from 20 million a quarter earlier. Anyone maintaining separate Copilot licences, training or internal documentation should expect the surface to consolidate.
Read more →Executive Order 14409 gave Treasury, the NSA and CISA until 1 August to deliver a classified process for benchmarking how far an AI model can find and exploit software weaknesses on its own, plus a voluntary framework giving the government 30 days to review a new frontier model before release. The NSA Director holds final say on which models count as covered. It is voluntary on paper, but labs that skip it will be conspicuous.
Read more →From 2 August, generative AI systems with more than a million monthly users in California must embed provenance data in AI-generated images, video and audio, offer a free detection tool, and pass the same obligations to licensees. Text is exempt. Penalties run to $5,000 per violation, with each day counted separately, and the date was chosen to line up with the EU transparency deadline.
Read more →Anthropic reported that an unreleased Mythos preview model, running semi-autonomously inside a Claude Code scaffold, derived a key-recovery attack on the post-quantum signature scheme HAWK that cut the cost of breaking its smallest key from 2^64 to 2^38 operations. It also produced a 200- to 800-fold speedup against seven-round AES-128. Neither result breaks anything in production, but two years of expert review had missed the HAWK weakness.
Read more →A survey of 2,400 knowledge workers by Writer and Workplace Intelligence found 29% had actively worked against their employer's AI plans, rising to 44% among Gen Z respondents. A separate poll of 1,005 US workers put active resistance to new AI tools at half the workforce. Rollouts that never answer what staff personally get out of the change tend to find out this way.
Read more →A short read every week — the few things worth your time, and nothing that isn’t.