{"id":240023,"date":"2025-10-21T02:47:19","date_gmt":"2025-10-21T02:47:19","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/240023\/"},"modified":"2025-10-21T02:47:19","modified_gmt":"2025-10-21T02:47:19","slug":"breaking-news-deepseek-opens-source-again","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/240023\/","title":{"rendered":"Breaking News: DeepSeek Opens Source Again"},"content":{"rendered":"<p>A picture is worth a thousand words! The DeepSeek-OCR model boldly explores the boundary of visual-text compression. By decoding more than 10 times the text information from a small number of visual tokens, this end-to-end VLM architecture not only outperforms GOT-OCR2.0 on the OmniDocBench benchmark but also provides an efficient solution to the long-context problem of LLMs.<\/p>\n<p>DeepSeek has released a new model!<\/p>\n<p>On Github, DeepSeek has created a new repository for DeepSeek-OCR, aiming to explore the boundary of visual-text compression.<\/p>\n<p>As the saying goes, a picture is worth ten thousand words. This is also true for LLMs!<\/p>\n<p>Theoretically, the DeepSeek-OCR model has preliminarily verified the feasibility of &#8220;context optical compression&#8221; \u2014<\/p>\n<p>The model can effectively decode more than 10 times the number of text tokens from a small number of visual tokens.<\/p>\n<p>That is to say, a single image containing document text can represent rich information with far fewer tokens than the equivalent text.<\/p>\n<p>This indicates that optical compression through visual tokens can achieve a higher compression ratio.<\/p>\n<p>As an intermediate modality connecting vision and language, the OCR task is an ideal testbed for the visual-text compression paradigm \u2014<\/p>\n<p>It establishes a natural compression-decompression mapping relationship between visual and text representations and provides quantifiable evaluation metrics.<\/p>\n<p>DeepSeek-OCR has high practical value in the OCR task: In the OmniDocBench benchmark test, it surpasses GOT-OCR2.0 (256 tokens per page) with only 100 visual tokens; with fewer than 800 visual tokens, it outperforms MinerU2.0 (more than 6000 tokens per page on average).<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,470\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014836_716_interlace,1.jpeg\"\/><\/p>\n<p class=\"img-desc\">Figure (a) shows the compression ratio (the number of real text tokens \/ the number of visual tokens used by the model) in the Fox benchmark test; Figure (b) shows the performance comparison on the OmniDocBench.<\/p>\n<p>In practical applications, a single A100-40G graphics card can support the generation of more than 200,000 pages of training data for large language models\/visual language models per day.<\/p>\n<p>The new model can also parse charts, chemical equations, simple geometric figures, and natural images:<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,1313\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014836_511_interlace,1.jpeg\"\/><\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,1330\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014836_770_interlace,1.jpeg\"\/><\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,1321\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014837_867_interlace,1.jpeg\"\/><\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,1243\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014837_277_interlace,1.jpeg\"\/><\/p>\n<p>In different historical context stages, the visual-text compression of DeepSeek-OCR can reduce the number of tokens by 7 &#8211; 20 times, providing a feasible direction for solving the long-context problem of large language models.<\/p>\n<p>This paradigm opens up new possibilities for rethinking the collaborative fusion of visual and language modalities and further improving the computational efficiency of large-scale text processing and intelligent agent systems.<\/p>\n<p>This discovery will strongly promote the future development of visual language models and large language models.<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,651\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014837_434_interlace,1.jpeg\"\/><\/p>\n<p>Github: https:\/\/github.com\/deepseek-ai\/DeepSeek-OCR<\/p>\n<p>HuggingFace: https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-OCR<\/p>\n<p>  The open-source artifact DeepSeek-OCR explores context optical compression<\/p>\n<p>Currently, open-source VLMs (visual language models) adopt three main visual encoder architectures, but each has its own drawbacks.<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,291\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014838_308_interlace,1.jpeg\"\/><\/p>\n<p>With the progress of VLMs, many end-to-end OCR models have emerged, fundamentally changing the traditional pipeline architecture and simplifying the OCR system.<\/p>\n<p>But there is a core question:<\/p>\n<p>For a document containing 1000 characters, at least how many visual tokens are needed for decoding?<\/p>\n<p>This question is of great significance for studying the principle of &#8220;a picture is worth a thousand words&#8221;.<\/p>\n<p>DeepSeek-OCR aims to answer this question. It adopts a unified end-to-end VLM architecture, consisting of an encoder and a decoder.<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,314\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014838_542_interlace,1.jpeg\"\/><\/p>\n<p>The encoder (i.e., DeepEncoder) is responsible for extracting image features, tokenizing and compressing the visual representation. The decoder generates the required results based on the image tokens and prompt information.<\/p>\n<p>  Encoder: The innovative architecture of DeepEncoder<\/p>\n<p>To verify the feasibility of &#8220;context optical compression&#8221;, the visual encoder needs to meet the following characteristics:<\/p>\n<p>   Be able to process high-resolution images;<br \/>\n   Maintain low activation overhead at high resolutions;<br \/>\n   Generate fewer visual tokens;<br \/>\n   Support multi-resolution input;<br \/>\n   Have a moderate parameter scale.<\/p>\n<p>The researchers proposed a brand &#8211; new visual encoder, DeepEncoder. DeepEncoder has approximately 380 million parameters and is mainly composed of SAM &#8211; base and CLIP &#8211; large connected in series.<\/p>\n<p>The visual perception feature extractor mainly uses window attention, and its main architecture is SAM &#8211; base (patch &#8211; size 16) with 80 million parameters;<\/p>\n<p>The visual knowledge feature extractor uses dense global attention, and its main architecture is CLIP &#8211; large with 300 million parameters.<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"874,408\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014838_12_interlace,1.jpeg\"\/><\/p>\n<p>Between these two components is a 2 &#8211; layer convolutional module that performs 16\u00d7 downsampling on the visual tokens.<\/p>\n<p>DeepEncoder will compress the image size. For example, an image with an input size of 1024\u00d71024 is divided into 1024\/16\u00d71024\/16 = 4096 patch tokens.<\/p>\n<p>The first half of the encoder is dominated by window attention and has only 80M parameters, so the activation memory consumption is acceptable.<\/p>\n<p>Before entering the global attention module, the 4096 tokens pass through the compression module, and the final number of tokens will be reduced to 4096\/16 = 256, making the overall activation memory consumption controllable.<\/p>\n<p>Suppose there is an image containing 1000 optical characters. To test how many visual tokens are needed for decoding, the model needs to support a variable number of visual tokens.<\/p>\n<p>That is to say, DeepEncoder needs to support multiple resolutions.<\/p>\n<p>Dynamic interpolation position encoding can meet the above requirements.<\/p>\n<p>The researchers designed multiple resolution modes to support multiple resolutions simultaneously during the model training process, thus enabling a single DeepSeek &#8211; OCR model to support multiple resolutions.<\/p>\n<p>As shown in Figure 4 below, DeepEncoder mainly supports two input modes: native resolution and dynamic resolution. Each mode contains multiple sub &#8211; modes.<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,460\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014838_903_interlace,1.jpeg\"\/><\/p>\n<p>The native resolution supports four sub &#8211; modes: Tiny, Small, Base, and Large.<\/p>\n<p>The dynamic resolution is composed of two native resolutions.<\/p>\n<p>Supporting dynamic resolution is mainly to meet the application requirements of ultra &#8211; high &#8211; resolution input (such as newspaper images). Tiling is a secondary window attention method that can further effectively reduce the activation memory consumption.<\/p>\n<p>In Gundam mode, the number of visual tokens output by DeepEncoder is n\u00d7100 + 256, where n is the number of tiles.<\/p>\n<p>Gundam mode is trained together with the four native resolution modes to achieve the goal of a single model supporting multiple resolutions.<\/p>\n<p>It is worth noting that the Gundam &#8211; master mode (a local view of 1024\u00d71024 + a global view of 1280\u00d71280) is obtained by continuing to train on the already trained DeepSeek &#8211; OCR model.<\/p>\n<p>Table 1 below summarizes the resolutions and the number of tokens in each mode.<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,330\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014839_484_interlace,1.jpeg\"\/><\/p>\n<p>  Decoder: DeepSeek &#8211; 3B &#8211; MoE<\/p>\n<p>The decoder uses DeepSeekMoE, specifically DeepSeek &#8211; 3B &#8211; MoE.<\/p>\n<p>During the inference process, the model activates 6 routing experts and 2 shared experts, with a total of approximately 570 million parameters activated.<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,702\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014839_590_interlace,1.jpeg\"\/><\/p>\n<p>The 3B DeepSeekMoE is very suitable for domain &#8211; centered visual language model (VLM) research \u2014<\/p>\n<p>It can obtain the expressive ability of a 3B model while enjoying the inference efficiency similar to that of a 500M small &#8211; scale model.<\/p>\n<p>  Specific results<\/p>\n<p>On the Fox benchmark set, the researchers verified the compression and decompression ability of DeepSeek &#8211; OCR on text &#8211; dense documents and preliminarily explored the feasibility and boundary of &#8220;context optical compression&#8221;.<\/p>\n<p>As shown in Table 2 below, within a 10\u00d7 compression ratio, the decoding accuracy of the model can reach approximately 97%, which is very promising.<\/p>\n<p>Moreover, the output format is not exactly the same as that of the Fox benchmark, so the actual performance may be slightly higher than the test result.<\/p>\n<p class=\"image-wrapper\"><img decoding=\"async\" data-img-size-val=\"1080,525\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/10\/1761014839_837_interlace,1.jpeg\"\/><\/p>\n","protected":false},"excerpt":{"rendered":"A picture is worth a thousand words! The DeepSeek-OCR model boldly explores the boundary of visual-text compression. By&hellip;\n","protected":false},"author":2,"featured_media":240024,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[4,450,132459,132460,132461,132457,132463,451,12736,132464,3,132462,452,453,132465,132458],"class_list":["post-240023","post","type-post","status-publish","format-standard","has-post-thumbnail","category-breaking-news","tag-breaking-news","tag-breakingnews","tag-context-optical-compression","tag-deepencoder","tag-deepseek-3b-moe","tag-deepseek-ocr","tag-fox-benchmark","tag-headlines","tag-large-language-models","tag-long-context-problem","tag-news","tag-omnidocbench","tag-top-stories","tag-topstories","tag-visual-language-models","tag-visual-text-compression"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/240023","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=240023"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/240023\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/240024"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=240023"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=240023"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=240023"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}