Brief IA

Meta: Muse Manages Permissions and Activities, Warp Announces PR in 35 Minutes

🛠️ AI Tools·Tom Levy·

Meta: Muse Manages Permissions and Activities, Warp Announces PR in 35 Minutes

Meta: Muse Manages Permissions and Activities, Warp Announces PR in 35 Minutes
Key Takeaways
1Warp claims to generate a pull request in 35 minutes, with the first human review occurring 3.5 hours later.
2Muse provides an activity feed detailing tools, scripts, research, and steps for each task.
3Muse failed to purchase New Balance 9060 in a specific color via the browser.
💡Why it mattersThese demonstrations confront a public agent and a software factory on concrete processes, highlighting execution metrics and design choices around permissions and transparency.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

A user test details how Muse, Meta's personal agent, manages appointments, family briefings, and goals while finely calibrating permissions. Meanwhile, Warp showcases its "factories" capable of transforming a Slack request into a pull request, which it claims is ready in 35 minutes, with a human review arriving hours later. Two visions of AI, consumer-oriented and engineering-focused, converge on transparency, browser limitations, and very concrete metrics.

At Warp, human review arrives 3.5 hours after a PR generated in 35 minutes

Warp indicates that the humans responsible for code review have now become the main bottleneck in its engineering process. According to the company, the first human review occurs on average 3.5 hours after the creation of a pull request, while the Wilson factory takes 35 minutes to go from launch to PR. On the consumer side, Claire believes that the use of the browser remains the weak point of agents, citing the example where Muse failed to purchase specific colored New Balance 9060s, returning incorrect results and not completing the purchase.

Contextual permissions and detailed activity log at the heart of Muse

Claire explains that Muse requests permission precisely when access is necessary, validates the obtained information, and seeks additional confirmation before taking any action. The agent does not immediately require total access and does not bombard the user with repeated approval requests. Claire believes that the way permission is requested is as important as what the agent can do once access is granted. Muse provides an activity feed that logs tool calls, scripts, searches, and individual steps. Claire thinks that most consumers will ignore this feed, but developers and advanced users will find great value in it. During her test, Muse neither exposed her to a terminal, nor confronted her with confusing errors, nor solicited her at the wrong time.

From Slack to merge: Wilson handles Linear tickets, QA, and PR without a keyboard

Zach Lloyd demonstrates that Warp Factories transforms a request sent on Slack into a pull request that will then be tracked, tested, and merged. The Wilson factory manages the entire process: triage, opening Linear tickets, development, generating pull requests, quality assurance (QA), and merging, all without any keyboard input required from a human. Warp states that a first pull request can be available in 35 minutes and that its factories handle 2,000 pull requests each month. For home use, Muse generated Claire's best family briefing on the first attempt, in the form of a more polished and structured PDF, complete with a "Talk at the Table" section and proactive detection of scheduling conflicts. For personal goals, Muse gathered context from multiple exchanges, created a reminder, and treated the goal as a continuous follow-up.

Measuring and improving: AI judge, task replay, and Pareto at Warp

Warp proposes tracking human interactions per pull request as an indicator of a factory's efficiency and presents the factory as covering the entire development cycle in the cloud. According to the company, the most relevant benchmark model comes from the organization's previous achievements rather than public rankings. It repeatedly executes actual factory tasks with various model configurations, subjects them to the same evaluation method used in production, and positions the trade-offs between cost and quality on a Pareto chart. Warp also specifies that factories improve when their configuration is integrated into the code: a learning loop gathers failed executions until a recurring pattern emerges. A scoring layer and an AI judge evaluate each execution across multiple dimensions, including the creation of redundant tests.

Consumer-friendly design of Muse and parallel business uses at Warp

Claire believes that Muse's originality for a non-technical audience stems from the design experience gained at Meta. The agent favors accessible language and avoids specialized terms like crons, tools, artifacts, plugins, and connectors, structuring the interface around a feed, ideas, goals, and a library. She also mentions an animated avatar designed for the AI, which uses a small computer to perform tasks and a sphere to generate media. In the business context, Zach Lloyd specifies that AI adds more value outside of engineering when it accesses concrete business data, and he demonstrates in real-time three tasks being performed simultaneously to identify customer questions and priorities regarding the website and commercial support. Warp also emphasizes that Wilson tasks are carried out in a shared Slack channel, allowing junior collaborators to learn by observing seasoned users, without the need for a structured training program.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.