⚡
Brief IA
›

Vibe Coding: Quick Gains, Security Risks

💻 Code & Dev·Tom Levy·

Vibe Coding: Quick Gains, Security Risks

Vibe Coding: Quick Gains, Security Risks
⚡
Key Takeaways
1SWE-Agent with Claude 4 Sonnet: 61% correct solutions, 10.5% secure
2METR study: +19% time with AI tools on mature repositories
3Controlled experiment: HTTP JavaScript task 55.8% faster with Copilot
💡Why it matters — Gains on bounded tasks coexist with security risks and reduced efficiency in mature codebases, necessitating precise safeguards.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

AI-assisted coding accelerates prototypes and certain repetitive tasks, achieving speeds up to 55.8% faster in a controlled experiment. However, recent measurements reveal blind spots: only 10.5% of secure solutions in a reported case for a coding agent, and developers were 19% slower with AI tools on mature repositories. Adoption is progressing nonetheless, including among Y Combinator startups, while raising widespread skepticism about the accuracy of outputs.

Security and Reliability Under Scrutiny

Security emerges as a significant weakness in code produced by agents. A benchmark observed that solutions could be functionally correct while remaining insecure. In a reported result, SWE-Agent with Claude 4 Sonnet delivered correct solutions 61% of the time, but only 10.5% of them were secure. An application may thus accomplish the requested task while exposing it to vulnerabilities. For anything related to authentication, payments, private data, permissions, uploads, APIs, or production databases, ignoring a security review turns vibe coding into a liability.

Beyond security, the confusion between "functional" and "correct" looms. Robustness is not merely about the happy path where forms, pages, graphics, and redirects seem to work, but in the unhappy cases: malformed files, expired APIs, inadequate permissions, null fields, concurrent edits, or hostile inputs. Vibe coding can deliver something that appears finished before it has been satisfactorily validated; a running demo does not equate to truly operational software.

In Mature Repositories, AI Can Slow Down by 19%

The notion that AI tools consistently speed up developers is not confirmed everywhere. A randomized controlled trial conducted by METR, involving experienced open-source maintainers working on their own mature repositories, shows that when these tools are allowed, the time to completion increases by 19%. Participants expected a gain; the study found the opposite.

These observations are found in the context of advanced codebases, which are marked by project history, established conventions, tacit architecture, testing, review standards, non-obvious dependencies, and maintainability requirements. In such contexts, AI tends to shift the workload from writing to the stages of review, correction, and integration. Vibe coding is better suited for small systems and visible requirements, losing efficiency on large, legacy, domain-specific codebases that are harder to contextualize for a model.

Prototypes and Bounded Tasks: Where It Really Speeds Up

The best terrain for vibe coding remains prototyping: moving from an idea to a prototype faster, whereas setting up a frontend, backend, database, authentication, deployment, environment variables, and error handling once formed a barrier to entry. AI coding tools lower this threshold, allowing for the description of a workflow and quickly obtaining a testable first draft. Usage has spread within the startup ecosystem; a quarter of Y Combinator's W25 cohort reportedly reached codebases generated 95% by AI, which does not guarantee quality but indicates real use in serious workflows.

For repetitive and well-defined tasks, the gains are tangible. A controlled experiment showed that a JavaScript HTTP server task could be completed 55.8% faster with GitHub Copilot, but the complexity of mature production limits the generalization of this result across all professions. For transformation scripts, API adapters, form validations, test skeletons, simple interface components, SQL queries, or configuration files, AI leverages the presence of frequent patterns. When these deliverables undergo testing, review, and present minimal risks, imposing manual line-by-line writing becomes more of a principle than a productivity imperative.

From Intention to Iteration: A Different Way to Code

Vibe coding shifts the entry point: instead of first choosing a framework, files, API connections, or a component structure, the user starts with intention and specification. Requesting a dashboard that ingests a CSV, filters clients based on churn risk, and exports a segment becomes the unit of work, with syntax following. The more precise the specification, the more faithful the result.

This approach thrives in a short loop of "ask, execute, observe, correct, test, repeat." It allows for reacting to an interface, errors, and real behaviors rather than documentation, provided it remains grounded in reality. The most effective users are those who demand proof of value at each iteration.

Opening Creation Without Losing Judgment

By making development more conversational, vibe coding opens creation to diverse profiles: analysts, product managers, researchers, marketers, founders, and domain experts. It becomes possible for an analyst to launch a Streamlit app, for a teacher to generate quizzes, for a marketer to track a campaign, for a researcher to tool annotation, or for a founder to prototype before hiring. This broadening is welcomed, while acknowledging a risk: producing useful tools without spontaneously integrating "production-ready" criteria. Tolerable for small uses, this lack becomes perilous when it touches on money, health, privacy, security, or vital processes.

The other pitfall lies in the relationship to learning. Vibe coding can make one productive sooner, but it can also short-circuit the acquisition of engineering judgment if one iterates blindly until "it works" without understanding the changes, creating dependence rather than competence. Developers are aware of this: in 2025, a Stack Overflow survey reported a distrust of the accuracy of AI tools (46%), exceeding trust (33%), with only 3% expressing total confidence. Best practices involve examining diffs, asking for explanations, writing tests, and checking edges, treating AI as a collaborator rather than an authority.

Ultimately, the main quality issues stem primarily from unarticulated elements. Requesting the creation of a login page will indeed generate a screen, but there is no guarantee that password hashing, limiting the number of attempts, session expiration, email verification, CSRF protection, cookie security, account locking, audit logging, or OAuth reminder management will be considered. Experienced teams mitigate these unspoken requirements by formulating precise demands: prohibiting CSVs larger than 10 MB, checking for empty files, preventing API key storage on the frontend, writing unit tests for incorrect inputs, using environment variables, implementing role-based access control, and logging import failures.

In total, software creation is becoming faster, more conversational, and more accessible, without making a running application a reliable system. The future is neither about replacing engineers with AI nor rejecting vibe coding; it will be nuanced. Those who know how to specify, test, review, and understand will move faster; those who refuse will primarily deliver fragilities. Vibe coding is not a substitute for judgment: it is a multiplier, for better or worse.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.