Can AI Surpass Textbook Authors?

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
AI Writing: Between Promises and Disappointing Realities
The author of this article recently wrote a manual on artificial intelligence and poses a crucial question: how long will it take before AI can perform this task better than he can? AI-generated writing has sparked numerous criticisms, especially when it comes to creative or voice-driven texts, such as blogs. These critiques, including those from many experts, emphasize that true writing is characterized by a distinct voice, a personal perspective, and a profound expression of humanity. These elements are often absent from the outputs of large language models (LLMs), which, although becoming more sophisticated tools, seem to regress in their ability to produce truly inspiring writing.
In contrast, non-fiction writing, which includes explanatory and informative texts, initially appeared to be a domain where LLMs could excel. Sam Altman, in a recent podcast, mentioned the ability of LLMs to generate filler text and useful copy. However, shortcomings in this area are often attributed to a lack of general intelligence in the models or training issues. Despite improvements, these models struggle to achieve a high level of non-fiction writing quality.
The Challenges of Non-Fiction Writing for LLMs
The stagnation of models in producing long-form non-fiction writing is concerning, especially for those who hope these technologies can autonomously solve major scientific problems. Currently, models struggle to organize and present even the most established sciences convincingly. This poses a significant barrier to their ability to independently tackle complex and open-ended problems. As long as these challenges remain unaddressed, the progress of LLMs in the scientific domain will be limited to solving simple problems and making connections between different fields, rather than achieving groundbreaking discoveries.
This position may seem paradoxical to those who are optimistic about AI advancements, especially at a time when companies like Anthropic announce progress on famous mathematical problems. For example, Anthropic published a blog post about Claude, which made strides on the famous Riemann Hypothesis. However, the scope of scientific problems is vast, and it is unlikely that current AI models can cover them as broadly as some believe.
Why is Progress Slow?
The author expected much more progress in non-fiction writing from AI models. He almost felt naïve for planning to publish a non-fiction book in the future, given the current state of affairs. Some of the most renowned models in writing capability, such as OpenAI's GPT 4.5 and Moonshot's Kimi K2, are already considered outdated. Around these releases, models have transitioned from acceptable to human-level performance in other tasks like coding and mathematics. However, writing well seems to be an orthogonal skill to most of these. The author does not believe that writing is simply being overlooked, but rather that it is difficult and lacks good training data for specific intervention.
There are certainly easy opportunities to improve AI models in writing—such as specialized harnesses like Claude Code, prompts, and training environments that make models spend significantly more inference tokens on output—but the author does not think this will have a multiplicative impact on their capability. Writing well is a very challenging task! It is unfortunate that we have not unlocked the inference scale for one of the great intellectual quests.
Current Limitations of AI Models
Today, models seem really poor at long-form technical writing. They can produce a correct sentence, but if you try to get them to write an entire chapter, it will be a mix of confusing formulations, poorly organized, and generally a bit off. They try to be too clever where it is unnecessary and, in the process, make random conceptual errors. Models will improve significantly on small mistakes in the near future, especially as they grow larger—allowing them to retain more knowledge—but the author does not expect their ability to utilize that knowledge to transform.
For example, GPT models have been incredible at finding typos and minor issues for a long time. The author passed an almost-final draft of his book in PDF format to GPT 5.5 Pro, and it found surprising minor typos in the 200-300 page manuscript. In contrast, Claude models have been much more helpful as editors. They have much better taste, tend to understand the mental model of the task better, and offer more interesting suggestions to unblock various forms of writer's block.
The examples the author provided above all share a consistent theme. The models know how to check each unit of content, usually a sentence, equation, or figure, or create a specific section where you are stuck. With these skills, they do not do a good job revisiting components and linking them together, as they pile many things on top of each other. This resembles a kind of irreducible errors that accumulate. We used to deal with these errors in mathematics and code, but upon reflection, RLVR has been a true magic solution for reducing them.
Extracting Value from Current Models as a Writer
The author is willing to share that there are a few phrases of technical explanation in his book that come from an AI model—well under 1%—they are there because he genuinely appreciated them. He allowed himself to consider including a few AI tokens in the book, as it did not feel like cheating if, as a true expert, he thought the phrase was what the reader needed. Especially in the editing process, where he had a very careful eye on things and many concerns about whether his book would ever be finished with everything he had to manage, this was an extremely valuable path. For instance, he had a list of questions from his editor interwoven in a LaTeX file with a specific delimiter like \editor{}. He let Claude Code navigate to each comment, print the context before and after, and let him know if it was a simple typo correction or something more nuanced. He would write a response—the text to insert—or ask Claude for suggestions before making corrections. Intellectually, this is a very focused editing process; it was a fun way to enhance the book. Sometimes, phrases from Claude's suggestions found their way into the book.
It is definitely a slippery slope, and when the author accepted a few AI suggestions, it was at the moment he was realizing his second complete revision of the manuscript. Emotionally, the project felt finished, but he still had work to do. Upon exiting the process of writing the manual, he greatly appreciates the strict rule he has for his writing on Interconnects to never use AI outputs in the content. It is much more enjoyable to write in a way that is solely you—voice-driven, so highly valued for the process—but writing a standard manual is not really an activity known for being fun. He understands why people turn AI tools into crutches when most of their writing is just a product to fill space, rather than a means to achieve a goal. He is motivated to write abundantly to learn, feel, and express.
He is also working on similar balances in his scientific work. AI models are excellent for repetitive parts of the paper, such as drafting a related works or context section that he knows by heart, but using them for the abstract, introduction, experiments, or conclusion is a shame. That is where the story and soul of the work are communicated—where you learn what your research truly means.
The author is convinced that he has created much more net value by being able to have AI models generate and verify his non-fiction writing. They make writing equations trivial, can help refactor the repository, translate between languages, and much more. At first, it was very enjoyable until he became a bit exhausted by the length of the publication process, watching the field advance.
As an example of why AI was crucial in this case, he had to maintain simultaneous Markdown and LaTeX versions of his book in two places, while readers were providing feedback on the web version and his editorial team at Manning was reviewing a derived copy. Without AI agents, synchronizing between the two would have easily taken him five times longer (and that task has already taken dozens of hours).
Something intertwined with this story, which the author discovered while reflecting on agents, is that your pace of understanding will not accelerate by using agents. This understanding, in the form of intuition, taste, instinct, etc., is what will be valuable in the future. Using AI for non-fiction writing takes away from that progression. Moreover, if you were not already an expert, you would not be able to grasp its flaws.
In his case, he felt such an urgency to pour the knowledge from his mind onto the page that there were moments when using AI models was a valuable tool. A significant part of the motivation for his book was to have a unique reference for important post-training methods like rejection sampling or character training, where there are very few resources.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.