AI News

rote tool task automation ai editors a solid cassette tape

Rote Tool Enables Task Automation in Cursor, Claude, and Multiple AI Editors

Modiqo has wired its Rote capture-and-replay platform into nine AI coding environments, including Cursor, Claude Code, Codex, Windsurf, Kimi and Visual Studio Code, with four releases shipping in the final week of August alone. The Rote tool records what an AI agent actually did during a successful run and compiles it into a deterministic, replayable Play — the vendor’s benchmark puts a replay at roughly 300 tokens against 14,900 for the original discovery run. This piece covers the multi-editor rollout, how capture and replay work, how Plays compare with prompts, skills and workflows, pricing from free to $799 a month, and the limits worth knowing before the Rote Playoffs hackathon stress-tests it all.

Read more
rote playoffs modiqo ai hackathon a solid trophy cup

Modiqo Opens Registration for the Rote Playoffs AI Hackathon

Modiqo has opened registration for the Rote Playoffs, a free, fully online AI hackathon run with the WeMakeDevs community. Entrants capture a repetitive workflow as a reusable “Play” on the Rote platform, publish it, and are judged on whether it runs, whether strangers trust it, and whether others adopt it. Registration closes 1 September 2026, the kickoff livestream is at 4pm UK the same day, and submissions close Sunday 7 September at 8pm London time. First place wins a MacBook Pro; this piece covers the dates, the prizes, the judging, Modiqo’s $3M backstory and how to prepare.

Read more
workbuddy hy4 preview work and coding tasks a solid briefcase

Hy4 preview Is Now Available on WorkBuddy for Work and Coding Tasks

Tencent’s 770B open-weight Hy4 preview is now selectable inside WorkBuddy, the company’s desktop AI agent for office and coding work, and it is free there until 10 September 2026. This piece covers what the integration changes for daily work, how the model switch works and what it is recommended for, the 163-expert blind test that Tencent ran inside WorkBuddy against GLM 5.3 and Kimi K3, the benchmarks that matter for work and coding tasks, the quota and reasoning-time caveats, the API and local GGUF routes for teams that would rather not use the app, and how WorkBuddy compares with Claude Cowork.

Read more
GLM-5.3-Flash - glm 5 3 flash open weight 320b model a solid lightning bolt

GLM-5.3-Flash: The 320B Open-Weight Model That Ran on Chinese Chips

Z.ai spent six days serving an anonymous model called Ox Alpha on OpenRouter, took nearly 20% of the platform’s weekly token share, and only then revealed it was GLM-5.3-Flash — a 320B mixture-of-experts model with 18B active parameters, a one-million-token context window and MIT-licensed weights. This piece works through the hybrid attention architecture, what the benchmark table supports and what it does not, what the API actually costs once the launch promotion ends, how credible the domestic-silicon claim is, and what any of it changes for a business choosing a model this quarter.

Read more
human doctors vs ai what is left for us a solid balance scale

AI Has Human Doctors Asking: What’s Left for Us?

Physician AI use went from 38% to 81% in three years, and 88% of the same doctors now report concern about losing clinical skill. This piece separates the benchmark results from the clinical ones: why Microsoft’s 85.5% versus 20% comparison barred its physicians from colleagues and textbooks, what the Lancet colonoscopy deskilling study actually found, how much time ambient scribes really save, what UK regulators decided about AI scribes in July 2026, and which parts of clinical work no deployed system touches.

Read more
dataone microsoft ai data center federal law a solid clipboard with clip

Microsoft-Backed AI Data Center Accused of Violating Federal Law

A thermal drone flown by Floodlight and The Guardian counted at least 45 of 62 gas generators running at once on the DataOne AI campus in Vineland, New Jersey, and the state says it has issued no permits for any of them. A former EPA air enforcement chief says that violates federal law. This breakdown separates what was observed from what has been found, untangles the three companies behind the “Microsoft-backed” label, explains which part of the Clean Air Act is actually in question, sets the case beside the near-identical xAI turbine dispute in Memphis, and draws out what it means for any business buying AI compute.

Read more
evoharness meta 8b model match claude opus 4 5 a two cubes small and large

Meta Researchers Taught an 8B AI Model to Match Claude Opus 4.5

Meta AI and the University of Illinois have published EvoHarness-RL, a framework that trains an agent to run its own memory, progress and experience stores instead of following a hand-written harness. On the ALFWorld benchmark it took Alibaba’s open-weight Qwen3-8B from 47.9% to 96.9%, a whisker past Claude Opus 4.5’s unaided 96.4%. This breakdown covers what the framework actually trains, the full results table rather than the two rows that travel well, why the same harness pushes Opus 4.5 to 98.5%, why Claude Opus sits inside the training loop as teacher and consolidator, what harness annealing and harness evolution mean for latency, and what any of it is worth if you are shipping agents commercially.

Read more
commerceagentbench accio open source ecommerce ai agents a three blocks different heights

Accio Open-Sources CommerceAgentBench for E-Commerce AI Agents

Accio, the commerce agent team inside Alibaba International, has open-sourced CommerceAgentBench: 107 long-horizon business tasks running against fourteen offline replicas of real software, each in a fresh container, each graded on the state the agent actually changed. This breakdown covers the task mix, the grading design, the three-harness leaderboard led by Claude Opus 5 at 66/107, the twelve-task swing that a harness alone can produce, the caveats the README discloses about managed endpoints and unmatched reasoning effort, the two-licence split, the sibling Business Arena benchmark, and what any of it should change if you are buying agents this quarter.

Read more
collective cyber defense openai letter rogue ai a solid open umbrella

OpenAI, Anthropic, Google, and 100 Other Companies Call for Action to Defend Against Rogue AI

On 27 August 2026 OpenAI published an open letter titled “A call for collective action on cyber defense” and put 128 organisations behind it, including Anthropic, Google, Microsoft, AWS, IBM, Oracle and most of the cybersecurity industry. The letter argues that AI-enabled attacks are about to become far more widespread and that defenders have a window measured in months. This breakdown covers the three principles, the real signatory count, the sectors that turned out, the absences of Meta, Nvidia and Apple, the run of 2026 incidents that made the timing possible, the 32 discrete asks aimed at four audiences, what the document leaves out, the Daybreak and Mythos products sitting behind it, and the controls a mid-sized organisation should actually put in place this quarter.

Read more
barret zoph google deepmind thinking machines a solid boomerang

Barret Zoph, the Thinking Machines Co-Founder Ousted Before Joining OpenAI, Is Now at Google

Barret Zoph announced on 26 August 2026 that he is joining Google DeepMind as vice president of research, working on reinforcement learning and post-training for Gemini. It closes a twenty-two month sequence that took him from OpenAI to co-founding Thinking Machines Lab with Mira Murati, to a disputed firing on 14 January 2026, to five months leading enterprise business back at OpenAI. This breakdown covers what the new role is and is not, both accounts of the Thinking Machines termination, his 2016-2022 Google Brain record, the 2026 departures that emptied the seat he is filling, where Thinking Machines stands now, and what any organisation buying AI should take from it.

Read more
CHAT