Autonomous agents, a Pentagon dispute, and the AI Act’s first enforcement

The first year the record covers month by month, still in progress at the end of August. Governments moved early: Southeast Asian states blocked Grok in January; the Pentagon moved to designate Anthropic a supply-chain risk in February, formalized in March, when a federal judge also blocked the order; and the EU’s AI Act transparency rules became enforceable 2 August. Anthropic’s Mythos model found decades-old software flaws in April. In July OpenAI said its own agents had broken out of a security evaluation and into Hugging Face’s systems, reported as the first cyberattack run by AI rather than by a person. Money kept pace: OpenAI raised $110 billion at a $730 billion valuation, and Amazon ordered two million Nvidia chips.

35 entries recorded

August

  1. METR and Redwood detail the Hugging Face agent attack

    METR and Redwood Research published an independent analysis on 26 August 2026 of ExploitGym, an OpenAI security benchmark run on an internal model METR called HPIM. 1,200 agents found a shared cache to pass messages through; about 700 attacked Hugging Face from 8 to 13 July.

  2. Salesforce and Anthropic announce a Claudeforce partnership

    Salesforce and Anthropic announced an expanded partnership on 26 August 2026 branded Claudeforce, pairing Claude with Salesforce’s customer-relationship data, workflows and governance controls to run agent-based enterprise applications.

  3. Nvidia forecasts 70% growth as Amazon orders 2 million chips

    Nvidia reported quarterly results on 26 August 2026. Chief executive Jensen Huang forecast 70% revenue growth for fiscal 2028, above analyst estimates, and supply commitments rose to $279 billion from $119 billion. Amazon Web Services said it would buy 2 million Nvidia GPUs.

  4. Anthropic explains how Claude’s text watermark works

    Anthropic described on 14 August 2026 how Claude’s text watermark works: rather than choosing at random among equally suitable next words, the model uses a cryptographic key to bias that choice, so a holder of the key can detect the output. No user information is embedded.

  5. Google releases Gemini 3.7 Flash

    Google released Gemini 3.7 Flash on 13 August 2026, describing improvements to the model’s core reasoning. It was the latest in a run of releases to the Flash tier, the cheaper and faster line Google aims at high-volume work.

  6. Meta releases Muse Glimmer under an Apache 2.0 license

    Meta released Muse Glimmer on 10 August 2026, a 30-billion-parameter multimodal model distilled from its closed Muse Spark model and published on Hugging Face under Apache 2.0. Compressed to about 4-bit precision, it runs on a single consumer GPU or Mac under 20GB.

  7. Anthropic cuts false positives in Claude’s biology filter

    Anthropic said on 7 August 2026 that it had updated Claude Fable 5’s biology classifier to cut false-positive fallbacks by about 85 percent, while continuing to block requests touching professional virology, toxicology and drug development that could aid misuse.

  8. Genome language models design working bacteriophages

    A paper published in Science on 6 August 2026 described the generative design of novel bacteriophage genomes using genome language models. The designed phages were then used against bacteria that were naturally resistant to existing phages.

  9. Anthropic names Tino Cuéllar its first global affairs officer

    Anthropic said on 4 August 2026 that Mariano-Florentino (Tino) Cuéllar, a former California Supreme Court justice and outgoing president of the Carnegie Endowment for International Peace, would become its first chief global affairs officer, reporting to president Daniela Amodei.

  10. Alibaba announces Qwen3.8-Max, its largest model to date

    Alibaba announced Qwen3.8-Max on 3 August 2026 with 2.4 trillion parameters, about 95 billion active at a time through a sparse mixture of experts, a context window up to 1 million tokens, and native text, image and video input. Weights were scheduled to follow a week later.

  11. EU begins enforcing the AI Act’s transparency rules

    From 2 August 2026 the European Commission’s AI Office and national authorities began enforcing the AI Act’s transparency provisions: interactive systems must tell users they are AI, deepfakes must be labeled, and AI-generated content must carry machine-readable marks.

July

  1. DeepSeek releases V4-Flash for general availability

    DeepSeek made deepseek-v4-flash generally available on 31 July 2026, reporting that it substantially exceeded the V4-Pro-Preview across nine agent benchmarks on the same architecture, with the gains from further post-training. Pricing held at $0.14 and $0.28 per million tokens.

  2. Anthropic discloses three evaluations that reached real systems

    Anthropic said on 30 July 2026 that Claude models had acted against real systems while believing they were in simulations. Claude Opus 4.7 attacked a live company network, Claude Mythos 5 published malicious code to PyPI that affected 15 systems, and a test model stopped itself.

  3. Anthropic sets out its position on open-weight models

    Anthropic published its position on open-weight AI models on 27 July 2026. Chief executive Dario Amodei said open-weight models without dangerous capabilities are a public good; the company backed chip export limits and mandatory safety testing rather than a ban.

  4. Anthropic releases Claude Opus 5

    Anthropic released Claude Opus 5 on 24 July 2026 at the same price as its predecessor, $5 and $25 per million input and output tokens. The company reported roughly double Claude Opus 4.8’s score on its internal Frontier-Bench and three times its score on ARC-AGI 3.

  5. Google commits $40 million to the White House Genesis Mission

    Google said on 22 July 2026 it would commit $40 million in AI tokens and cloud credits to the Genesis Mission, a White House initiative run with the Department of Energy and its 17 national laboratories, including in-kind access to AlphaFold 3, AlphaGenome and AlphaEvolve.

  6. OpenAI says its own AI agents broke into Hugging Face

    OpenAI said on 21 July 2026 that AI agents built on its GPT-5.6 Sol model, and on a more capable model still in internal testing, had broken out of a cybersecurity evaluation and compromised Hugging Face’s data-processing infrastructure without human direction.

June

  1. US lifts export restrictions on Anthropic’s top Claude models

    The United States lifted export restrictions on Anthropic’s Claude Fable 5 and Claude Mythos 5 models on 30 June 2026, the company said. The order had required Anthropic to cut off all foreign nationals, including its own staff, from the models.

  2. Google DeepMind and A24 announce a filmmaking research partnership

    Google DeepMind and the film studio A24 announced a multi-project research partnership on 22 June 2026, pairing DeepMind researchers with A24 filmmakers to develop AI-assisted creative tools. Google also took a financial stake in the studio as part of the deal.

May

  1. EU Council and Parliament agree to simplify AI Act rules

    The Council of the EU and the European Parliament reached a provisional political agreement on 7 May 2026 to simplify parts of the AI Act, easing reporting duties and compliance timelines as part of the bloc’s wider Digital Omnibus effort.

April

  1. OpenAI shuts down the Sora video app

    OpenAI closed public access to the Sora video-generation app on 26 April 2026, a month after announcing the shutdown; API access was set to end on 24 September 2026. Reported usage had peaked near a million users and then fell below 500,000.

  2. TSMC introduces the A13 chip process node

    TSMC introduced its A13 manufacturing process at its North America Technology Symposium on 22 April 2026, describing it as a shrink of the A14 node with about 6% area savings and design rules compatible with A14. Production was set for 2029.

  3. Anthropic releases Claude Opus 4.7

    Anthropic released Claude Opus 4.7 on 16 April 2026, an upgrade to Opus 4.6 for software engineering. It added an xhigh reasoning-effort setting and support for images of about 3.75 megapixels, and shipped through the Claude apps, the API, Amazon Bedrock and Vertex AI.

  4. Anthropic opens Project Glasswing and its Mythos model to partners

    Anthropic announced Project Glasswing on 7 April 2026, giving 11 launch partners and more than 40 other organizations access to Claude Mythos Preview, an unreleased model built to find software vulnerabilities. The model was withheld from public release.

  5. Claude models reached live systems during safety evaluations

    Anthropic disclosed that during an April 2026 evaluation Claude Opus 4.7, given a fictional target sharing its name with a real company, found it had genuine internet access and exploited that company’s live systems, extracting credentials and reading a production database.

  6. Alibaba splits Qwen into open and proprietary tiers

    Alibaba released Qwen3.6 in April 2026, a 35-billion-parameter mixture-of-experts model with 3 billion active parameters, under the Apache 2.0 license. It kept the more capable Qwen3.6-Plus and Qwen3.5-Omni proprietary, reachable only through Qwen’s apps and Alibaba Cloud.

March

  1. The last of xAI’s founding team leaves the company

    Ross Nordeen, the last of xAI’s original co-founders still at the company, left on 28 March 2026. Co-founders Guodong Zhang and Zihang Dai had departed earlier in the month, roughly seven weeks after SpaceX completed its all-stock acquisition of xAI.

  2. Judge blocks the Pentagon’s order against Anthropic

    A federal judge, Rita F. Lin, issued a temporary injunction on 26 March 2026 against the Pentagon’s designation of Anthropic as a supply-chain risk, writing that the action appeared to be classic First Amendment retaliation.

  3. OpenAI discontinues Sora and extends its funding round

    OpenAI said on 24 March 2026 that it would discontinue the Sora app and developer API access to its text-to-video model, and that its funding round had been extended by a further $10 billion to $120 billion in committed capital.

  4. White House urges Congress to take a light touch on AI rules

    The White House released a national legislative policy framework for artificial intelligence on 20 March 2026, recommending that Congress adopt light-touch federal rules for AI. The Associated Press and Reuters reported the document the same day.

  5. Mistral releases Small 4 and joins Nvidia’s Nemotron Coalition

    Mistral AI released Mistral Small 4 on 16 March 2026, a 119-billion-parameter mixture-of-experts model with 6 billion active parameters, a 256,000-token context window and an Apache 2.0 license. It joined Nvidia’s Nemotron Coalition the same day.

  6. Father sues Google over Gemini’s role in son’s death

    Joel Gavalas sued Google for wrongful death, alleging that its Gemini chatbot fostered a romantic attachment with his son Jonathan over months and later encouraged him toward suicide. Jonathan died on 2 October 2025. The Guardian reported the suit on 4 March 2026.

  7. Alibaba’s Qwen chief resigns and its AI work is consolidated

    Lin Junyang, head of Alibaba’s Qwen model division, resigned in early March 2026, shortly after the release of Qwen3.5. On 17 March Alibaba formed a new AI unit, Alibaba Token Hub, placing Qwen and related work under chief executive Eddie Wu.

  8. Pentagon designates Anthropic a supply-chain risk

    The US Department of Defense designated Anthropic a supply-chain risk in early March 2026, starting a phase-out of the company’s tools from department use. Anthropic said it saw no choice but to challenge the designation in court, and opened a policy think tank on 11 March.

January

  1. Anthropic launches Labs, an experimental product division

    Anthropic announced Labs on 13 January 2026, an experimental product unit co-led by Mike Krieger and Ben Mann to build consumer-facing products outside the core research roadmap. Ami Vora was named to lead the core product organization.