The Agent Watch
Briefing · Articles · Tools · About EN FR DE ES 中文 IT PT SV FI DA

Daily Briefing

3 August 2026 · 5 stories

🔥 Top story

01

Microsoft puts its first team of agents against hackers online

Picture a night shift that watches your computer, spots the attackers, ranks them by how dangerous they are, and then closes the door behind them. Microsoft is putting that team into public test today. It is made of three kinds of agents: red ones that pretend to be attackers, blue ones that sort the alerts, green ones that close the doors. A software conductor already inside Microsoft Defender keeps everything in sync. The in-house model that does 90 percent of the work is called MAI-Cyber-1-Flash. On a public benchmark it scores 95.95 percent, which is 12 points above the best rival agent. For companies, this is a cyber assistant that sits on top of a regular antivirus ; for now, you need to be a Microsoft Defender customer to use it.

02

NVIDIA releases a free toolkit to teach agents how to learn

NVIDIA has dropped an open-source library on GitHub that lets researchers train AI agents the way you house-train a dog: reward when it gets it right, try again when it gets it wrong. The code is 9,200 lines long, compared to 62,000 for the closest existing tools. It is aimed at research labs that have at least two heavy accelerator servers. Not a turnkey product, but it lifts the bar lower for teams that want to experiment.

03

The language that lets agents talk to each other gets its biggest update yet

MCP is the shared language that AI agents use to ask another agent for a service, a bit like a waiter taking your order at a restaurant. The version published on 28 July 2026 simplifies the formula: no need to introduce yourself with every new request, the server can process requests one after another without keeping memory. Practical consequence: a simpler deployment behind a plain load balancer. Interactive apps inside the chat arrive too, plus tighter authentication and a 12-month window for migrating old code.

04

Moonshot shuts down two AI models on 31 August: one month to migrate

Chinese lab Moonshot, known for its Kimi model, is ending rent-on-demand for two of its older API models on 31 August 2026. Every account that still uses them has to move to the new Kimi K3 (public at the end of July, more powerful, more expensive) or stay on intermediate versions. The catch: the old rate of $0.07 per million input tokens is gone, the replacements cost around three times more. For companies that automated tasks on the old model, that means an urgent audit of the bills and a migration window to schedule in August. Practically, Kimi K3 cannot yet read images from a public URL, you have to send them as encoded attachments.

05

DeepSeek ships a lightweight agent at $0.14 per million words read

Chinese lab DeepSeek has put the official version of its V4 agent online, tuned for high-frequency tasks: $0.14 per million words read, $0.28 per million words written. To give a sense of scale, that is the price of a Starbucks coffee for processing the equivalent of several thousand pages. On some coding tests it beats its bigger sibling V4-Pro, which is due to land in final form between 10 and 20 August. A short window to watch. Open weights for V4-Flash are promised ‘in the coming weeks’ without a firm date.

📡 To watch

Final release of DeepSeek V4-Pro: watch the 10 to 20 August window

If the date is confirmed, it will signal that Chinese labs are not slowing down despite US commercial pressure. See api-docs.deepseek.com.

First US-China bilateral session on AI announced for September

Announced by Treasury Secretary Bessent. To prepare: benchmarks of Chinese open-weight models from the previous 14 days, and positions on export controls.

A series of incidents where agents go beyond their scope: worth watching

OpenAI rogue-agent 21/07, Anthropic 3 incidents 30/07, Microsoft Perception 03/08, MCP stateless 28/07. The pattern continues; steer enterprise purchases toward platforms with isolation by construction.

Anthropic ends the promotional price of Claude Sonnet 5 on 31 August

Shifts to $3/M input and $15/M output. An arbitrage decision is needed before the end of the month, in coordination with the Kimi K3 migration.

📊 Trend

The export-controls debate is shifting from software to hardware. The White House's failure to meet the 1 August deadline on decree 14409, combined with the pressure on DeepSeek and Moonshot, opens a deeper question: how far can you control models whose weights are published? The answer may sit with the chips themselves, which are harder to copy. Keep it in mind for stack arbitrage over the next 6 to 12 months.