⚡
Brief IA
›

3D Scenes from a Photo: GPT-6 Astra Shines but Misjudges Itself

🛠️ AI Tools·Tom Levy·

3D Scenes from a Photo: GPT-6 Astra Shines but Misjudges Itself

3D Scenes from a Photo: GPT-6 Astra Shines but Misjudges Itself
⚡
Key Takeaways
1Coding agents generate 3D Blender scenes from a photo
2GPT-6 Astra achieves 53.4% indoors and 39.6% outdoors on LEGO-Bench
3The self-assessment of agents is unreliable; a plugin based on objective metrics mainly improves weaker models
4The scenes are used for computer vision but remain less accurate than specialized models
💡Why it matters — This work demonstrates the potential of coding agents for 3D reconstruction while highlighting the need for objective measures to progress towards reliable applications.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Researchers from Maryland and AWS demonstrate that an agent can code a 3D Blender scene from a single image. Tested on a new benchmark, the models often deliver usable results but struggle to evaluate themselves and reproduce geometry. A plugin based on objective metrics significantly improves the weakest variants, while GPT-6 Astra achieves the best scores, with measured ceilings.

A Persistent Gap and an Ecosystem Already in Motion

A significant gap remains between a functional scene and an accurate reconstruction. Researchers believe that executable scene programs generated by coding agents show potential but are not yet precise enough. The superiority of GPT-6 Astra measured on LEGO-Bench aligns with other findings. According to AI researcher Yoav Artzi, this model represents a significant advancement in spatial understanding and was likely trained on large volumes of 3D data, such as Blender scenes. 3D software companies are already preparing for the arrival of these agents: Unity now offers official plugins for Claude Code and Codex. Other approaches reconstruct scenes directly within the model, such as Atlas from World Labs. Google DeepMind is taking a different path with GenCeption, a video model for depth estimation and segmentation that achieves the performance of specialized models.

What the Scores on LEGO-Bench Show

During testing, the six GPT models evaluated almost always generated a functional scene, but accuracy varied widely. GPT-6 Astra scored 53.4% on indoor scenes and 39.6% on outdoor scenes, while the weaker configurations reached around 15%. GPT-6 Astra proved to be the closest to the reference images. Older models often omitted camera angles, lighting, or entire objects. Accuracy decreases with the complexity of the scenes, and outdoor scenes are more challenging than indoor ones. By increasing the reasoning budget, the variants of GPT-6 improved significantly: on a subset of office tests, Astra's score rose from 32.3% to 61.8%. Initial attempts at work are often mediocre, revisions can undo previous progress, and self-evaluation remains unreliable.

Self-Evaluation Fails; a Plugin Enforces Metrics and Boosts the Weak

When it comes to choosing between two versions of a scene, the geometric judgments of the models are close to or below the level of chance. An agent struggles to reliably assess whether its scene has improved. Researchers believe that refinement should rely on tangible metrics rather than the agent's appraisal. They developed LEGO-Plugin, an extension that operates without requiring additional training phases, which pairs the initial scene with the reference image, substitutes self-evaluation with objective criteria, and preserves correct advancements against modifications that would degrade the scene. The plugin improved the performance of all six tested models, with particularly significant gains for the weaker ones, achieving up to a 62.7% increase, while the already high-performing model gained about two percentage points.

The LEGO-Bench and Its Scoring Method

LEGO-Bench serves as an evaluation tool and comprises 208 images from 104 scenes, with 443 recorded assets. Researchers emphasize that obtaining an accurate 3D truth from real photos is challenging and that simple synthetic scenes lack realism. To reconcile these requirements, images are rendered from professionally designed simulators, with geometry, depth, and object assignments remaining hidden to allow for automated scoring. The complexity of the scenes can be increased without altering lighting or camera settings. Evaluation is conducted along three axes: Validity, which checks that a usable artifact has been produced; Reconstruction, which measures visible geometric accuracy; and Appearance, which compares a pixel-by-pixel re-render to the reference image.

LEGO-Anything Generates a Blender Script from a Photo

LEGO-Anything, developed by researchers from Maryland and AWS, tasks an agent with producing a Blender program from an image using the Image-to-Code method. The agent progresses through successive steps: it writes, executes, inspects, and then adjusts the code, with the output explicitly describing the objects, geometry, layout, and camera position. This executable format can be controlled and modified like computer code, but researchers note weaknesses in self-evaluation and geometric fidelity. The reconstructed scenes have served as a basis for computer vision tasks such as object detection, segmentation, and depth estimation, extracted directly from the program. Without additional training, the results are usable but remain modest: object detection achieves about half the performance of the specialized model DINO, while the gap is larger for segmentation and depth compared to references like SAM 3 and Depth Anything 3.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.