3D Scenes from a Photo: GPT-6 Astra Shines but Misjudges Itself

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Researchers from Maryland and AWS demonstrate that an agent can code a 3D Blender scene from a single image. Tested on a new benchmark, the models often deliver usable results but struggle to evaluate themselves and reproduce geometry. A plugin based on objective metrics significantly improves the weakest variants, while GPT-6 Astra achieves the best scores, with measured ceilings.
A Persistent Gap and an Ecosystem Already in Motion
A significant gap remains between a functional scene and an accurate reconstruction. Researchers believe that executable scene programs generated by coding agents show potential but are not yet precise enough. The superiority of GPT-6 Astra measured on LEGO-Bench aligns with other findings. According to AI researcher Yoav Artzi, this model represents a significant advancement in spatial understanding and was likely trained on large volumes of 3D data, such as Blender scenes. 3D software companies are already preparing for the arrival of these agents: Unity now offers official plugins for Claude Code and Codex. Other approaches reconstruct scenes directly within the model, such as Atlas from World Labs. Google DeepMind is taking a different path with GenCeption, a video model for depth estimation and segmentation that achieves the performance of specialized models.
What the Scores on LEGO-Bench Show
During testing, the six GPT models evaluated almost always generated a functional scene, but accuracy varied widely. GPT-6 Astra scored 53.4% on indoor scenes and 39.6% on outdoor scenes, while the weaker configurations reached around 15%. GPT-6 Astra proved to be the closest to the reference images. Older models often omitted camera angles, lighting, or entire objects. Accuracy decreases with the complexity of the scenes, and outdoor scenes are more challenging than indoor ones. By increasing the reasoning budget, the variants of GPT-6 improved significantly: on a subset of office tests, Astra's score rose from 32.3% to 61.8%. Initial attempts at work are often mediocre, revisions can undo previous progress, and self-evaluation remains unreliable.
Self-Evaluation Fails; a Plugin Enforces Metrics and Boosts the Weak
When it comes to choosing between two versions of a scene, the geometric judgments of the models are close to or below the level of chance. An agent struggles to reliably assess whether its scene has improved. Researchers believe that refinement should rely on tangible metrics rather than the agent's appraisal. They developed LEGO-Plugin, an extension that operates without requiring additional training phases, which pairs the initial scene with the reference image, substitutes self-evaluation with objective criteria, and preserves correct advancements against modifications that would degrade the scene. The plugin improved the performance of all six tested models, with particularly significant gains for the weaker ones, achieving up to a 62.7% increase, while the already high-performing model gained about two percentage points.
The LEGO-Bench and Its Scoring Method
LEGO-Bench serves as an evaluation tool and comprises 208 images from 104 scenes, with 443 recorded assets. Researchers emphasize that obtaining an accurate 3D truth from real photos is challenging and that simple synthetic scenes lack realism. To reconcile these requirements, images are rendered from professionally designed simulators, with geometry, depth, and object assignments remaining hidden to allow for automated scoring. The complexity of the scenes can be increased without altering lighting or camera settings. Evaluation is conducted along three axes: Validity, which checks that a usable artifact has been produced; Reconstruction, which measures visible geometric accuracy; and Appearance, which compares a pixel-by-pixel re-render to the reference image.
LEGO-Anything Generates a Blender Script from a Photo
LEGO-Anything, developed by researchers from Maryland and AWS, tasks an agent with producing a Blender program from an image using the Image-to-Code method. The agent progresses through successive steps: it writes, executes, inspects, and then adjusts the code, with the output explicitly describing the objects, geometry, layout, and camera position. This executable format can be controlled and modified like computer code, but researchers note weaknesses in self-evaluation and geometric fidelity. The reconstructed scenes have served as a basis for computer vision tasks such as object detection, segmentation, and depth estimation, extracted directly from the program. Without additional training, the results are usable but remain modest: object detection achieves about half the performance of the specialized model DINO, while the gap is larger for segmentation and depth compared to references like SAM 3 and Depth Anything 3.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.