SmolDocling - The SmolOCR Solution?
Key Takeaways
SmolDocling, an OCR solution, is compared to other open and proprietary OCR solutions, with a focus on its capabilities and applications, utilizing tools such as Hugging Face and GitHub.
Full Transcript
Okay, so a new week has come and yet another OCR model is here. So this time this is kind of an interesting model that comes out of hugging face and I think they've partnered with IBM on this. So this is small dockling. So if you've seen Huging Place has made a whole bunch of what they call small models generally these are really tiny sort of around the 1B or under kind of size and it's basically a document understanding model not just purely OCR and it's only 256 million parameters. So, the advantage here that this has got is that it can run on GPUs that don't have a lot of VRAM, but in my early playing around with it, it still seems to me that you actually still do need a GPU to actually get this working. So, it's interesting here that they talk not just about OCR, but around the whole idea of document conversion, and they claim that this is beating every competing model we tested up to 27x. And that to me sounded really impressive until I came in and looked at the actual models that they're testing. Don't include things like M LCR, don't include things like the Mistral OCR, etc. Obviously, you're not going to include other proprietary models like Gemini or the Open AI models, etc. So, it really does seem to me that this is an interesting model, especially when comparing to other small or tiny sort of VLMs for doing this particular task. and also interesting in that they've created this not just for doing OCR but perhaps in many ways for doing document conversion and document understanding. Okay, so the original dockling project is also a really interesting project in that this is again not just OCR but for doing extraction from documents and you can see that they're supporting a whole bunch of different documents from not just things like PDFs but Microsoft Word files, HTML, images, etc. And the idea here is that the small dockling is basically continuing that with just a small VLM type model in here. All right. So they've released a paper. So if we look at the architecture that they've got in here, we've got a standard sort of VLM architecture. And we can actually see that basically they've based this on the small VLM architecture. So I think that's basically they're using a sig lip vision encoder of 93 million parameters and then also the small LM model which is 135 million parameters. And then obviously that combined with their projection layers gets them up around this 256 million parameter size. So if we look at one of the diagrams in the paper, we can see that they're not just going for OCR, they're actually giving us locations and what they're calling this sort of dock tags format. And this is basically describing the elements of whether it's a text element, a picture, a table, code, etc. And then also actually where it is on the page and then doing the OCR out. So if you look in here, we can see that when we get lists of things, we're actually getting almost like an HTML structure coming out of this of where it's showing it. This is a list item, what the location is, and then we've got the OCR coming out of this. So, I do kind of feel this is something that you'd probably then take and you could even then feed it into another LLM to get it to tidy it up, etc. All right, so the model is up on Hugging Face if you want to try it out. They've also updated their post on the small VLMs to basically add this in here as well, which is something that they've been working on for a while. There a whole bunch of these small LLMs and VLMs in here. Now when we look at the model card they talk a bit more about the whole sort of dock tag stuff and show us that like okay you can actually get variety of different things out of the docs. So things like code recognition, formula recognition, tables, charts. So one of the things you can do is then process those things out and then use custom models to actually extract information out of those. So currently you can run it using the transformers library or you can run it using the VLM library for faster batch inference as well. And in here we can see like some of the instructions that it's actually being trained with to extract things like charts out formulas tables and of course to be able to extract out the text itself as well. Okay. So they've got a demo up where you can try these out. And if I come in here and let me just bring up the image that it's going to be processing, you can see this is the image that it's going to go through where it's basically got some normal text. It's then got the code blocks again. Then it's got some normal text and it's got some other sort of outputs. So we're looking to see how does it handle all of this code, I guess, as well as this. Okay. So, we can see that it starts off with the markdown output. And it looks pretty nice, right? The way it's done that. Looking at it, it's processed all of this as a code block. It's missed going back there. So, it didn't go back to get that one and sort of continue. It basically just went right on through there. All right, let's try another one. So, this time we've given it a whole chart and we're asking it to convert the chart. And just looking at the chart, we can see that there's a bunch of stuff in there. We can see that it's gone through. It's worked out. It's worked out a bunch of information, but I wouldn't say that it's necessarily any easier for reading in here. Let's try one more. Okay, here we can see that it's basically extracting out different elements of basically a logo or picture it's saying for the top bit there. But it is able to get out different sort of text by the looks of this. So this seems to be in French. And I kind of feel that the real advantage of a model like this is not going to be as a general OCR model. I think looking at playing with these demos is that the real advantage of something like this is going to be when you fine-tune it that for most sort of tasks that people end up doing. Usually if you've got a very specific kind of task that you wanted to do, you're going to find that the inputs are going to be reasonably similar. meaning that they're all going to be receipts. They're all going to be some kind of thing. And I kind of feel like if you put the effort in, it looks like we've run into some GPU errors there, but I kind of feel if you've put the effort in to making your own sort of labeled data set, you're able then to fine-tune this model to do really well at the tasks that you want it to do. So overall playing with it in the demos, trying it on a few examples of my own, it certainly is not a state-of-the-art OCR model in general. It's probably state-of-the-art for its size and for what it's doing. But I kind of feel like that's not the key point here. I think the real key point here is that this is a really interesting model at this idea of document extraction and document conversion. And I kind of feel that with its size being so small, it really lends itself to you fine-tuning it for your specific task in here. So, Hugging Face already has some scripts up for doing the finetunes of the small VLM. My guess is that these can be repurposed pretty simply to do this kind of specialized OCR task as well, as long as you put in the effort and make some of the data yourself. All right. Overall, I don't think this is necessarily going to replace M O OCR or things like the Mistral OCR, Gemini, etc. for these general OCR tasks. But I do think this can be very useful for making your own document conversion pipeline. And maybe that's something I'll look at in a future video. Anyway, if you've got any questions, as always, please put them in the comments below. I'd love to hear what you think. Have you tried it out yourself? What you're seeing that it works well for? What do you see that it's not working well for? Let me know in the comments below. And as always, if you found the video useful, please click like and subscribe. And I will talk to you in the next video.
Original Description
In this video I look at SmolDocling and how it compares to the other OCR solutions that are out there, both open and proprietary.
Blog: https://huggingface.co/blog/smolervlm#smoldocling
Paper: https://arxiv.org/pdf/2503.11576
HF Model: https://huggingface.co/ds4sd/SmolDocling-256M-preview
Demo: https://huggingface.co/spaces/ds4sd/SmolDocling-256M-Demo
For more tutorials on using LLMs and building agents, check out my Patreon
Patreon: https://www.patreon.com/SamWitteveen
Twitter: https://x.com/Sam_Witteveen
🕵️ Interested in building LLM Agents? Fill out the form below
Building LLM Agents Form: https://drp.li/dIMes
👨💻Github:
https://github.com/samwit/llm-tutorials
⏱️Time Stamps:
00:00 SmolDocling Tweet on X
01:35 Docling Github
02:05 SmolDocling Paper
03:26 SmolDocling Hugging Face
03:40 SmolVLM Blog
04:22 SmolDocling Special Instructions
04:29 SmolDocling Demo
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Sam Witteveen · Sam Witteveen · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
LangChain Basics Tutorial #1 - LLMs & PromptTemplates with Colab
Sam Witteveen
LangChain Basics Tutorial #2 Tools and Chains
Sam Witteveen
ChatGPT API Announcement & Code Walkthrough with LangChain
Sam Witteveen
Trying Out Flan 20B with UL2 - Working in Colab with 8Bit Inference
Sam Witteveen
LangChain - Conversations with Memory (explanation & code walkthrough)
Sam Witteveen
LangChain Chat with Flan20B
Sam Witteveen
LangChain - Using Hugging Face Models locally (code walkthrough)
Sam Witteveen
PAL : Program-aided Language Models with LangChain code
Sam Witteveen
Building a Summarization System with LangChain and GPT-3 - Part 1
Sam Witteveen
Building a Summarization System with LangChain and GPT-3 - Part 2
Sam Witteveen
Microsoft's Visual ChatGPT using LangChain
Sam Witteveen
Building a Summarization System with LangChain - Part 3 Using ChatGPT Turbo
Sam Witteveen
LangChain Agents - Joining Tools and Chains with Decisions
Sam Witteveen
Investigating Alpaca 7B - Finetuned LLaMa LLM
Sam Witteveen
Comparing LLMs with LangChain
Sam Witteveen
Running Alpaca7B in Colab
Sam Witteveen
How to finetune your own Alpaca 7B
Sam Witteveen
How to make a custom dataset like Alpaca7B
Sam Witteveen
Understanding Constitutional AI - the paper and key concepts
Sam Witteveen
Using Constitutional AI in LangChain
Sam Witteveen
Talking to Alpaca with LangChain - Creating an Alpaca Chatbot
Sam Witteveen
Text-to-video-synthesis with Diffusers and Colab
Sam Witteveen
Meet Dolly the new Alpaca model
Sam Witteveen
Checking out the Cerebras-GPT family of models
Sam Witteveen
A Step-by-Step Guide to Fine-Tuning Your Dolly Model (tutorial)
Sam Witteveen
Is GPT4All your new personal ChatGPT?
Sam Witteveen
Raven - RWKV-7B RNN's LLM Strikes Back
Sam Witteveen
Talk to your CSV & Excel with LangChain
Sam Witteveen
Vicuna - 90% of ChatGPT quality by using a new dataset?
Sam Witteveen
Koala Revealed: The ChatGPT Alternative You Need to Know! 🔍
Sam Witteveen
Running Koala for free in Colab. Your own personal ChatGPT? (tutorial)
Sam Witteveen
BabyAGI: Discover the Power of Task-Driven Autonomous Agents!
Sam Witteveen
Auto-GPT - How to Automate a Task Based AI with GPT-4
Sam Witteveen
Improve your BabyAGI with LangChain
Sam Witteveen
Generative Agents - Deep Dive and GPT-4 Recreation
Sam Witteveen
GPT4ALLv2: The Improvements and Drawbacks You Need to Know!
Sam Witteveen
Dolly 2.0 by Databricks: Open for Business but is it Ready to Impress!
Sam Witteveen
Red Pajama - Operation: Freeing LLaMA
Sam Witteveen
Investigating Open Assistant - Models, Datasets and Addons
Sam Witteveen
Investigating MiniGPT-4 - The Secret behind GPT-V?
Sam Witteveen
Stable LM 3B - The new tiny kid on the block.
Sam Witteveen
Bard can now code and put that code in Colab for you.
Sam Witteveen
Checking out Bark: a Text to Speech system by Suno AI
Sam Witteveen
Fine-tuning LLMs with PEFT and LoRA
Sam Witteveen
Master PDF Chat with LangChain - Your essential guide to queries on documents
Sam Witteveen
Using LangChain with DuckDuckGO Wikipedia & PythonREPL Tools
Sam Witteveen
Building Custom Tools and Agents with LangChain (gpt-3.5-turbo)
Sam Witteveen
StableVicuna: The New King of Open ChatGPTs?
Sam Witteveen
WizardLM: Evolving Instruction Datasets to Create a Better Model
Sam Witteveen
LaMini-LM - Mini Models Maxi Data!
Sam Witteveen
Finding the Best Free ChatGPT
Sam Witteveen
MPT-7B - The First Commercially Usable Fully Trained LLaMA Style Model
Sam Witteveen
LangChain Retrieval QA Over Multiple Files with ChromaDB
Sam Witteveen
LangChain Retrieval QA with Instructor Embeddings & ChromaDB for PDFs
Sam Witteveen
LangChain + Retrieval Local LLMs for Retrieval QA - No OpenAI!!!
Sam Witteveen
Transformers Agent - Is this Hugging Face's LangChain Competitor?
Sam Witteveen
StarCoder - The LLM to make you a coding star?
Sam Witteveen
Testing Starcoder for Reasoning with PAL
Sam Witteveen
The New Wizards - Unfiltered & Unaligned
Sam Witteveen
Camel + LangChain for Synthetic Data & Market Research
Sam Witteveen
More on: Reading ML Papers
View skill →Related Reads
📰
📰
📰
📰
A lightweight workflow for keeping up with AI conference papers
Dev.to · Daniel
Why CitedEvidence Believes Great Researchers Read Less Than You Think
Medium · AI
How to Write a Literature Review That Actually Argues Something
Medium · Machine Learning
I Built a Personal Paper Engine to Stop Losing Research Papers
Dev.to · Ethan
Chapters (7)
SmolDocling Tweet on X
1:35
Docling Github
2:05
SmolDocling Paper
3:26
SmolDocling Hugging Face
3:40
SmolVLM Blog
4:22
SmolDocling Special Instructions
4:29
SmolDocling Demo
🎓
Tutor Explanation
DeepCamp AI