Brief IA

Granite 4.0 3B Vision: A Revolution for Business Documents

🔬 Research·Tom Levy·

Granite 4.0 3B Vision: A Revolution for Business Documents

Granite 4.0 3B Vision: A Revolution for Business Documents
Key Takeaways
1Granite 4.0 3B Vision is a compact model designed to extract information from complex documents, such as tables and graphs.
2The model uses ChartNet, a multimodal dataset, to enhance the understanding of graphs with 1.7 million samples.
3Thanks to the DeepStack architecture, Granite 4.0 3B Vision efficiently handles visual and semantic details for optimal performance.
💡Why it mattersThis model could transform document management in businesses, making data analysis more accurate and faster.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Granite 4.0 3B Vision: A Breakthrough for Business Documents

Granite 4.0 3B Vision marks a significant advancement in the field of artificial intelligence applied to business documents. This compact vision-language model is specifically designed to reliably extract information from complex documents, including forms, tables, and structured visuals. It stands out for its ability to effectively handle varied data structures, making it particularly useful for companies looking to automate and optimize their document management processes.

Table Extraction and Graph Understanding

One of the key features of Granite 4.0 3B Vision is its ability to analyze complex tables, whether multi-row or multi-column, from document images. This capability allows for precise extraction of structured data, facilitating the integration of this information into data management systems. Furthermore, the model excels in converting graphs and figures into machine-readable formats, enabling automated and accurate interpretation of visual data.

Identification of Semantically Significant Key-Value Pairs

Granite 4.0 3B Vision is also designed to identify and anchor semantically significant key-value pairs in various document formats. This feature is essential for extracting structured information from diverse documents, allowing for deeper analysis and more informed decision-making.

Integration with Granite 4.0 Micro and Docling

The model is offered as a LoRA adapter, seamlessly integrating with Granite 4.0 Micro. This allows for flexibility between vision-language tasks and text-only solutions, facilitating integration into mixed pipelines. In tandem with Docling, Granite 4.0 3B Vision enhances the visual processing capabilities of documents, providing in-depth visual understanding that is crucial for many professional applications.

Construction and Innovations of Granite 4.0 3B Vision

The performance of Granite 4.0 3B Vision is based on three main pillars: a custom-designed graph understanding dataset, an innovative architecture named DeepStack, and a modular design that keeps the model practical for enterprise deployment.

ChartNet: A New Dimension in Graph Understanding

Graphs present a particular challenge for vision-language models, as their understanding requires joint reasoning about visual patterns, numerical data, and natural language. To address this gap, ChartNet has been developed. This large-scale multimodal dataset, which will be detailed in a paper for CVPR 2026, includes 1.7 million samples of diverse graphs. What makes ChartNet unique is that it uses a code-guided synthesis pipeline to generate samples covering 24 types of graphs and 6 plotting libraries.

A Code-Guided Synthesis Pipeline

Each sample in ChartNet includes five aligned components: the plotting code, the rendered image, the data table, a natural language summary, and question-answer pairs. This approach provides a comprehensive cross-view of graphs, allowing models to move from simple description to true understanding of the structured information they encode.

DeepStack: An Intelligent Approach to Visual Feature Injection

Unlike most vision-language models, Granite 4.0 3B Vision utilizes DeepStack injection. This method routes abstract visual features to earlier layers for semantic understanding, while high-resolution spatial details are processed in later layers. This ensures precise understanding of documents, both in content and layout, which is crucial for tasks such as table extraction, graph understanding, and analysis of semantically significant key-value pairs.

Modularity and Flexibility

Granite 4.0 3B Vision is offered as a LoRA adapter on top of Granite 4.0 Micro, rather than as a standalone model. In practice, this means that the same deployment can serve both multimodal and text-only workloads, automatically reverting to the base model when vision is not required. This modularity simplifies enterprise integration without sacrificing performance.

Performance Evaluation

Granite 4.0 3B Vision has been rigorously evaluated across several benchmarks to demonstrate its effectiveness and capabilities.

Performance on Graphs

On the ChartNet benchmark, Granite 4.0 3B Vision achieved the highest score on Chart2Summary with 86.4%, even surpassing larger models. It also ranked second on Chart2CSV with 62.1%, just behind Qwen3.5-9B, a model more than twice its size.

Table Extraction

The model has been tested on cropped tables and full-page documents, achieving high scores on benchmarks such as PubTablesV2 and OmniDocBench. Granite 4.0 3B Vision delivered top performance across these benchmarks, ranking first on PubTablesV2 for both cropped (92.1) and full pages (79.3), OmniDocBench (64.0), and TableVQA (88.1) among all evaluated models.

Extraction of Semantically Significant Key-Value Pairs

On the VAREX benchmark, Granite 4.0 3B Vision achieved an accuracy of 85.5% in zero-shot mode, demonstrating its ability to extract precise information from complex forms. VAREX is specifically designed to discriminate between small extraction models, encompassing 1,777 U.S. government forms ranging from simple layouts to complex nested and tabular structures.

Practical Use Cases

Granite 4.0 3B Vision offers practical solutions for various business needs, transforming how companies manage and analyze their documents.

Form Processing and Financial Analysis

The model can extract structured fields from documents such as invoices, forms, and receipts, or analyze financial reports to extract actionable data. By using Docling, it is possible to process graphs using chart2csv, chart2code, and tables using tables_json to convert them into structured, machine-readable data.

Research Document Intelligence

With Docling, Granite 4.0 3B Vision enables efficient management of academic documents, making visual content as accessible as text. It allows for OCR and layout analysis across dense academic PDFs, passing extracted figures to chart2summary and table crops to tables_html to make visual content discoverable alongside free text in a single pipeline.

Granite 4.0 3B Vision is now available on HuggingFace, published under the Apache 2.0 license. All technical details, training methodology, and benchmark results are available in the model's technical sheet. Users are encouraged to share their experiences and feedback in the community tab.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.