Granite 4.0 3B Vision: A Revolution for Business Documents
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Granite 4.0 3B Vision: A Breakthrough for Business Documents
Granite 4.0 3B Vision marks a significant advancement in the field of artificial intelligence applied to business documents. This compact vision-language model is specifically designed to reliably extract information from complex documents, including forms, tables, and structured visuals. It stands out for its ability to effectively handle varied data structures, making it particularly useful for companies looking to automate and optimize their document management processes.
Table Extraction and Graph Understanding
One of the key features of Granite 4.0 3B Vision is its ability to analyze complex tables, whether multi-row or multi-column, from document images. This capability allows for precise extraction of structured data, facilitating the integration of this information into data management systems. Furthermore, the model excels in converting graphs and figures into machine-readable formats, enabling automated and accurate interpretation of visual data.
Identification of Semantically Significant Key-Value Pairs
Granite 4.0 3B Vision is also designed to identify and anchor semantically significant key-value pairs in various document formats. This feature is essential for extracting structured information from diverse documents, allowing for deeper analysis and more informed decision-making.
Integration with Granite 4.0 Micro and Docling
The model is offered as a LoRA adapter, seamlessly integrating with Granite 4.0 Micro. This allows for flexibility between vision-language tasks and text-only solutions, facilitating integration into mixed pipelines. In tandem with Docling, Granite 4.0 3B Vision enhances the visual processing capabilities of documents, providing in-depth visual understanding that is crucial for many professional applications.
Construction and Innovations of Granite 4.0 3B Vision
The performance of Granite 4.0 3B Vision is based on three main pillars: a custom-designed graph understanding dataset, an innovative architecture named DeepStack, and a modular design that keeps the model practical for enterprise deployment.
ChartNet: A New Dimension in Graph Understanding
Graphs present a particular challenge for vision-language models, as their understanding requires joint reasoning about visual patterns, numerical data, and natural language. To address this gap, ChartNet has been developed. This large-scale multimodal dataset, which will be detailed in a paper for CVPR 2026, includes 1.7 million samples of diverse graphs. What makes ChartNet unique is that it uses a code-guided synthesis pipeline to generate samples covering 24 types of graphs and 6 plotting libraries.
A Code-Guided Synthesis Pipeline
Each sample in ChartNet includes five aligned components: the plotting code, the rendered image, the data table, a natural language summary, and question-answer pairs. This approach provides a comprehensive cross-view of graphs, allowing models to move from simple description to true understanding of the structured information they encode.
DeepStack: An Intelligent Approach to Visual Feature Injection
Unlike most vision-language models, Granite 4.0 3B Vision utilizes DeepStack injection. This method routes abstract visual features to earlier layers for semantic understanding, while high-resolution spatial details are processed in later layers. This ensures precise understanding of documents, both in content and layout, which is crucial for tasks such as table extraction, graph understanding, and analysis of semantically significant key-value pairs.
Modularity and Flexibility
Granite 4.0 3B Vision is offered as a LoRA adapter on top of Granite 4.0 Micro, rather than as a standalone model. In practice, this means that the same deployment can serve both multimodal and text-only workloads, automatically reverting to the base model when vision is not required. This modularity simplifies enterprise integration without sacrificing performance.
Performance Evaluation
Granite 4.0 3B Vision has been rigorously evaluated across several benchmarks to demonstrate its effectiveness and capabilities.
Performance on Graphs
On the ChartNet benchmark, Granite 4.0 3B Vision achieved the highest score on Chart2Summary with 86.4%, even surpassing larger models. It also ranked second on Chart2CSV with 62.1%, just behind Qwen3.5-9B, a model more than twice its size.
Table Extraction
The model has been tested on cropped tables and full-page documents, achieving high scores on benchmarks such as PubTablesV2 and OmniDocBench. Granite 4.0 3B Vision delivered top performance across these benchmarks, ranking first on PubTablesV2 for both cropped (92.1) and full pages (79.3), OmniDocBench (64.0), and TableVQA (88.1) among all evaluated models.
Extraction of Semantically Significant Key-Value Pairs
On the VAREX benchmark, Granite 4.0 3B Vision achieved an accuracy of 85.5% in zero-shot mode, demonstrating its ability to extract precise information from complex forms. VAREX is specifically designed to discriminate between small extraction models, encompassing 1,777 U.S. government forms ranging from simple layouts to complex nested and tabular structures.
Practical Use Cases
Granite 4.0 3B Vision offers practical solutions for various business needs, transforming how companies manage and analyze their documents.
Form Processing and Financial Analysis
The model can extract structured fields from documents such as invoices, forms, and receipts, or analyze financial reports to extract actionable data. By using Docling, it is possible to process graphs using chart2csv, chart2code, and tables using tables_json to convert them into structured, machine-readable data.
Research Document Intelligence
With Docling, Granite 4.0 3B Vision enables efficient management of academic documents, making visual content as accessible as text. It allows for OCR and layout analysis across dense academic PDFs, passing extracted figures to chart2summary and table crops to tables_html to make visual content discoverable alongside free text in a single pipeline.
Granite 4.0 3B Vision is now available on HuggingFace, published under the Apache 2.0 license. All technical details, training methodology, and benchmark results are available in the model's technical sheet. Users are encouraged to share their experiences and feedback in the community tab.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.