⚡
Brief IA
›

AI Agents: $20,000 on Devin, Grok 4.6 Tested

🛠️ AI Tools·Tom Levy·

AI Agents: $20,000 on Devin, Grok 4.6 Tested

AI Agents: $20,000 on Devin, Grok 4.6 Tested
⚡
Key Takeaways
1Ryan Carson claims to have spent $20,000 in one month on Devin and manages 10 to 15 threads, up to 40 PR/day.
2Claire ranks Grok 4.6 at the top of her blind tests alongside GPT-5.6 Sol, ahead of Sonnet 5 and Opus 5.
3Grok Bot offers useful multi-account connectors from day one, while Cursor Origin is considered not yet ready.
💡Why it matters — These insights describe concrete uses and limitations of AI agents and models, as well as the adoption and management conditions set by practitioners.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

A founder claims to have spent $20,000 on Devin in one month, managing 10 to 15 threads of agents and up to 40 pull requests per day. He describes his methods, from handwritten lists to the Watchdog workflow, and his approach to hiring in the age of agents. Meanwhile, Claire ranks Grok 4.6 at the top of her blind tests alongside GPT-5.6 Sol, while preferring Sonnet 5 for agent exchanges. She details what Grok Bot brings immediately with its multi-account connectors, the limitations of Origin, and the emerging coherence of the Cursor/xAI stack.

Devin in Production: $20,000, 10–15 Threads, and Up to 40 PRs/Day

Ryan Carson claims to have invested $20,000 in Devin over the course of a month. He says he manages 10 to 15 Devin threads simultaneously and ships up to 40 pull requests per day. His threads are categorized as P0, P1, P2, and Bugs, and he treats each one as a manager would with a direct report: clear objectives, set priorities, without micromanagement.

Management Methods: A Handwritten List and the Watchdog Workflow

Although he has eight screens, Ryan Carson continues to use a handwritten list for his weekly priorities. He explains that he deliberately chose this method to stay focused on three main tasks, even in the face of a constant influx of notifications from agents. He has also built an AI playbook that operates like a customer success team: Watchdog reviews each account, extracts activity and errors from Sentry and internal logs, isolates three key issues, and checks if a recent pull request has resolved them. This all feeds into a single Devin thread that he executes when he lacks visibility, allowing him to quickly get back up to speed.

More Output Doesn’t Make a Better Product: Caution and Listening

Ryan Carson and Claire express skepticism about allowing agents to work unconstrained overnight. They believe that frontier models can easily produce a large volume without truly understanding the real needs of customers. For them, product ideas, priorities, and market judgment must come from a human in contact with users. Ryan Carson attributes Untangle's product-market fit to a meeting with attorney Renee Bauer and attentive listening, rather than an increase in delivered code.

The Right Tool in the Right Place: Devin, Codex, and Claude as Complements

Ryan Carson indicates that most engineering tasks are handled by cloud agents, and he uses Devin to manage bugs, pull requests, investor updates, and customer triage. He turns to Codex for features that present significant visual complexity and when he wants to stay very close to implementation, a choice he claims is based on latency concerns, browser access, and the ability to inspect live. Beyond the codebase, Claire utilizes Devin for operations, custom quotes, and triage, while Ryan Carson centralizes investor updates through a reusable skill, with the shared idea of an agent that understands the entire codebase to solve various problems. On the design side, Claire explains that she used Claude Design to transform a Figma file into tokens and specifications, then entrusted the implementation to Codex, which she finds better for converting specifications into shared components; she considers Claude more efficient for producing portable specifications.

Recruiting in the Age of Agents: The Screen Recording Test

To evaluate candidates, Ryan Carson requests a video recording showing the creation of a feature in an existing application, with the entire screen displayed and without an introductory interview. He then provides Devin for their use and examines how they interact with the agent through the replay of their session. He claims to observe their thinking, building, and management of an agent under realistic conditions. He finds this method more revealing than a traditional behavioral interview.

Grok 4.6 Tops Claire's Tests, Grok Bot Useful from Day 1, Origin Still Lacking

Claire reports that Grok 4.6 has risen alongside GPT-5.6 Sol at the top of her own index, ahead of Sonnet 5 and Opus 5. She emphasizes her method: blind evaluations that she scores herself, giving 70% of the final weight to her judgment, without relying on third-party rankings. For agent exchanges, she continues to prefer Sonnet 5, which she finds closer to a solid collaborator; she believes Grok 4.6 is not yet competitive in this area. Grok Bot has been useful to her from day one thanks to its multi-account connectors, capable of aggregating her four email addresses and seven Slack spaces into a single bot, a capability she believes is still unmatched by Codex and Claude. She highlights a quick setup, a clean interface, and functioning built-in plugins, while noting that profiles seeking extensive customization might find it too polished; she points out that her more "chaotic" OpenClaw agents stimulate her more.

Regarding Cursor Origin, she praises a vision of a native alternative to GitHub designed for agents, with Bugbot, Cursor, and a tailored pull request flow, but believes that currently the offering remains a more appealing variant of GitHub with fewer functions. According to her, teams heavily reliant on GitHub Actions, code owners, and their automations will need a much stronger incentive to migrate, and she is not ready to replace GitHub with Origin.

She also observes an overall coherence between Grok Bot for knowledge workers, Origin for code hosting, Grok 4.6 as the default model, and the Cursor IDE already central for many developers. No element is perfect in isolation, but together, she sees a credible enterprise platform emerging; she adds that large organizations often prefer a single vendor and that Cursor is increasingly positioning itself in this regard.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.