Brief IA

LLM: Double Testing and New Routines in Project Management

🔬 Research·Tom Levy·

LLM: Double Testing and New Routines in Project Management

LLM: Double Testing and New Routines in Project Management
Key Takeaways
1The estimated time spent on testing has roughly doubled with the use of LLMs
2Programming takes up a smaller share, while testing becomes the main bottleneck
3Detailed planning, command/goal setting, and automated testing via Playwright MCP are suggested to optimize organization
💡Why it mattersThe shift in time allocation requires rethinking project management and adopting new tools to leverage LLMs.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

In LLM-boosted workflows, the estimated time spent on testing has roughly doubled, while programming is on the decline. To keep pace, a method is proposed: plan meticulously, trigger a completion check with the /goal command, and assign automated tests to agents via Playwright MCP.

Testing Becomes the Bottleneck: Increased Load and Declining Coding

The estimated time dedicated to testing has roughly doubled, indicating a significantly heavier validation load. With LLMs writing code, the share of testing has noticeably increased, while programming is declining. Testing is presented as the new bottleneck, and the priority now is to reduce its duration. This shift in workload also frees up time for other tasks but requires a different organization.

Automating Validation: Playwright MCP and Browser Agents

An operational setup involves making Playwright MCP available to all agents, including Claude Code and Codex. Agents can then launch local servers and run test scenarios in Chrome. This loop allows them to detect errors introduced by their own code and save a lot of time. The autonomy gained fits into a broader logic: reducing human interactions by enhancing automated execution.

Anticipating Tasks to Limit Agent Ambiguity

More time is now invested upfront to plan the work. Tasks often arise from a product feedback or bug message, or are part of an ongoing project, and then undergo detailed preparation. Before delegating to agents, ambiguities are clarified as much as possible. Otherwise, agents quickly return with questions deemed costly. This clarification is achieved by considering the expected scope or discussing with an LLM to bring out the gray areas.

Ensuring Completion: /goal as a Trigger for Verification

The active use of the /goal command is presented as a key lever. It initiates a completion check for the agent regarding the request and, if necessary, encourages them to continue until full correction. It has been reported that, recently, particularly with Opus 5, the absence of /goal often accompanies unfinished work. The command is seen as a quick fix while waiting for better solutions, interpreting these shortcomings as a form of laziness on the part of the agents.

Time Reallocation: Indicated Shares and New Priorities

Time shares are proposed: 70% for coding, 30% for interaction with agents, 10% for meetings — a level described as stable — 30% for testing, and an additional 30% for other tasks. The extra time unlocked is described as an estimate, available for training, creating more agents, or accomplishing more work. More broadly, the arrival of LLMs is linked to a profound evolution in project management practices and a need for new optimizations. The proposed techniques should be adapted, with an invitation to generalize them cautiously according to each application domain.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.