The Agent Watch
Briefing Articles Tools About EN FR DE ES 中文 IT PT SV FI DA

Latest briefing

July 14, 2026 · 3 stories (site) · 3 stories (base)

On July 14, 2026, AI agent infrastructure continues to mature at a sustained pace: Fujitsu puts a vertical multi-agent platform into production to rewrite legacy systems, two new research frameworks show that an agent can now automatically rewrite its own software layer and gain 60% performance, and Bespoke Labs raises $40 million to build realistic training environments for autonomous agents. Three industrial building blocks — platform, harness, environment — that consolidate simultaneously.

🔥 Top story

01

Fujitsu puts into production a team of AI agents to rewrite legacy computer systems

Replacing the software that runs banks or hospitals — written in the 1980s — usually takes years of projects and costs fortunes. On July 14, Fujitsu — Japan's largest digital services group with 100,000 employees and $23 billion in revenue — launched in Japan a turnkey service that does the work with a team of specialised AI agents, each with its own role: one translates code, another verifies, a third fixes in a loop. The service promises to cut large modernisation projects by about 40%. It is the first industrial deployment at scale of a vertical multi-agent platform by a global Tier-1 player — until now, this kind of tooling had stayed confined to demonstrations.

02

Two new frameworks let an AI agent rewrite its own software layer — and gain +60% performance

Today, a high-performing AI agent needs to be 'well harnessed': prompts, memory, tools, error handling. Self-Harness and HarnessX, two papers published in late June, show that an agent can now observe its own execution traces, detect where it fails, then automatically rewrite this software layer to improve itself. Result: an open-source model goes from 40.5% to 61.9% on the Terminal-Bench-2.0 benchmark — without touching the weights. In practice, an agent's performance is no longer just a question of model choice: it's also a question of scaffolding, and that scaffolding can now be optimised automatically. For a company that was hesitant to depend on a single model provider, this is a promising diversification path.

03

A startup raises $40 million to build virtual offices where autonomous agents train

AI models are progressing fast, but the agents that use them still often stumble when chaining a long task. Bespoke Labs' bet, a California startup that just raised $40 million, is simple: for an agent to become as reliable as a colleague, it must train in environments that resemble a real office — code, emails, tickets, microservices, conversations — not on textbook exercises. The company already provides tools used by Anthropic, OpenAI and Google DeepMind to evaluate their frontier models. If agents are ever to become full colleagues, this kind of training ground will likely be one of the missing pieces.

📡 To watch

Fujitsu — the 90-day test: if the -40% holds up, it's a new category of service offerings

Fujitsu's service targets 100,000 enterprise customers in Japan and more than 100 countries through its international network. The announced -40% time savings still need to be confirmed on 3-4 real projects within 90 days. If they hold, the big digital services groups (Accenture, Capgemini, NTT Data, TCS) will be forced to replicate the model within 6 months — or lose their pricing power on large modernisation projects.\n

Self-Harness and HarnessX confirm the industry's centre of gravity is shifting from model to scaffolding

When an agent can rewrite its own harness with a 60% gain on a benchmark, competitive advantage no longer lies only in the model, but also in the quality of orchestration and self-improvement. Self-Harness and HarnessX (Xiaomi team) are becoming a new market segment in their own right. For integrators and agentic platform vendors, it's the signal to position on 'harness engineering' now.\n

Bespoke Labs becomes one of the agent-training infrastructure providers to watch — Terminal-Bench is already a standard

Anthropic, OpenAI and Google DeepMind already use Terminal-Bench to evaluate their frontier models. With $40 million in funding led by 8VC and Wing VC, and the backing of Jeff Dean and Tristan Handy (CEO of dbt Labs), Bespoke Labs is becoming a likely building block of the agent-training ecosystem. For agent builders, it's a partner to identify quickly.\n

📊 Trend

On July 14, 2026, three converging movements show that AI agent infrastructure is maturing at great speed. Fujitsu — Japan's largest digital services group — launches in production a vertical multi-agent platform that promises to cut legacy modernisation projects by 40%. Two new research frameworks (Self-Harness and HarnessX, Xiaomi team) show that an agent can now automatically rewrite its own software layer and gain 60% performance on a benchmark, without touching the model. Bespoke Labs raises $40 million to build realistic training environments used by Anthropic, OpenAI and Google DeepMind. For those building with AI, three lessons emerge: (1) vertical multi-agent platforms move from demos to world-class offerings — the entry ticket for large projects is dropping, (2) an agent's performance is no longer only in the model: it's also in the quality of the scaffolding around it, and that scaffolding can now be optimised automatically, (3) agent training is becoming an infrastructure segment in its own right — as model training became yesterday.