{"id":465732,"date":"2026-05-29T07:49:12","date_gmt":"2026-05-29T07:49:12","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/465732\/"},"modified":"2026-05-29T07:49:12","modified_gmt":"2026-05-29T07:49:12","slug":"run-step-3-7-flash-on-nvidia-gpus-with-enterprise-ready-multimodal-ai","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/465732\/","title":{"rendered":"Run Step 3.7 Flash on NVIDIA GPUs with Enterprise-Ready Multimodal AI"},"content":{"rendered":"<p>AI\u00a0applications are\u00a0moving\u00a0beyond text generation to\u00a0multimodal\u00a0systems that can perceive, search, and reason\u00a0across images, documents,\u00a0video,\u00a0and language in real time\u2014turning\u00a0fragmented\u00a0information into actionable insights.\u00a0\u00a0<\/p>\n<p><a href=\"https:\/\/huggingface.co\/stepfun-ai\/Step-3.7-Flash\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">Step 3.7 Flash<\/a>, the latest from\u00a0StepFun,\u00a0brings these\u00a0capabilities\u00a0to production and enterprise-scale, available\u00a0on NVIDIA-accelerated infrastructure.\u00a0It\u00a0is\u00a0a\u00a0198B-parameter Mixture-of-Experts vision-language model,\u00a0with approximately 11B activated parameters per forward pass,\u00a0optimized\u00a0for agentic workflows that combine\u00a0perception, search, and multi-step reasoning at production scale.\u00a0<\/p>\n<p>With\u00a0native\u00a0image\u00a0and video input, three configurable reasoning levels\u2014low, medium, and high\u2014and a 256k context window,\u00a0it\u00a0is designed\u00a0for enterprise use cases such as financial analysis, concurrent coding\u00a0agents,\u00a0and\u00a0other\u00a0high-throughput multimodal use cases.\u00a0Developers can use\u00a0StepFun\u2019s\u00a0NVFP4-quantized checkpoint available\u00a0through\u00a0Hugging Face for\u00a0boosted inference\u00a0due to\u00a0reduced memory bandwidth and storage requirements.\u00a0<\/p>\n<p>ModelStep\u00a03.7 Flash\u00a0Total parameters\u00a0198B\u00a0Visual encoder parameters\u00a01.8B\u00a0Active parameters\u00a011B\u00a0Context length\u00a0256K\u00a0Experts\u00a0288 (8 active)\u00a0Table 1. Overview of the key Step\u00a03.7 Flash specs, such as parameter counts, context length, and\u00a0MoE\u00a0configuration<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"1194\" height=\"918\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/QuietHarborDiagram.webp\" alt=\"A diagram that shows how text and images are processed by the model through the vision encoder and core language model to provide text output.\" class=\"lazyload wp-image-117441\" style=\"aspect-ratio:1.3006871189393205;width:676px;height:auto\"  data-\/><\/p>\n<p>\t\tFigure 1.\u00a0A high-level diagram of the\u00a0Step\u00a03.7\u00a0Flash\u00a0components for text and vision processing<\/p>\n<p>Step\u00a03.7\u00a0Flash\u00a0can be deployed with\u00a0open source\u00a0frameworks such as\u00a0\u00a0<a href=\"https:\/\/docs.sglang.io\/cookbook\/autoregressive\/StepFun\/Step-3.7-Flash\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">SGLang<\/a>,\u00a0NVIDIA <a href=\"https:\/\/nvidia.github.io\/TensorRT-LLM\/models\/supported-models.html\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">TensorRT-LLM<\/a>,\u00a0and\u00a0<a href=\"https:\/\/docs.vllm.ai\/projects\/recipes\/en\/latest\/StepFun\/Step-3.7-Flash.html\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">vLLM<\/a>\u00a0to utilize kernels\u00a0optimized\u00a0for NVIDIA hardware.<\/p>\n<p>Build with NVIDIA endpoints\u00a0<a href=\"#build_with_nvidia_endpoints\u00a0\" aria-label=\"Scroll to Build with NVIDIA endpoints\u00a0 section\" class=\"heading-anchor-link\"><\/a><\/p>\n<p>Developers can use GPU-accelerated endpoints available through build.nvidia.com for\u00a0prototyping\u00a0and evaluating Step 3.7 Flash.\u00a0Test this out in\u00a0the\u00a0<a href=\"https:\/\/github.com\/NVIDIA\/GenerativeAIExamples\/tree\/main\/oss_tutorials\/Nemotron_Parse_StepFun_Document_Intelligence\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">demo notebook<\/a>, which uses Step 3.7 Flash and\u00a0<a href=\"https:\/\/build.nvidia.com\/nvidia\/nemotron-parse\" target=\"_self\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"internal\">NVIDIA Nemotron\u00a0Parse<\/a>. The\u00a0multi-step document\u00a0intelligence\u00a0pipeline\u00a0extracts structured insights\u00a0from large, complex documents\u00a0with bounding boxes\u00a0like financial reports, slide decks, and scientific papers,\u00a0including PDFs, and organizes the output.\u00a0<\/p>\n<p>Video 1. See how document intelligence pipelines extract usable data, then follow the workflow in a JupyterLab notebook<\/p>\n<p>Production-ready deployment with NVIDIA NIM\u00a0<a href=\"#production-ready_deployment_with_nvidia_nim\u00a0\" aria-label=\"Scroll to Production-ready deployment with NVIDIA NIM\u00a0 section\" class=\"heading-anchor-link\"><\/a><\/p>\n<p><a href=\"https:\/\/www.nvidia.com\/en-us\/ai\/\" target=\"_self\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"internal\">NVIDIA NIM<\/a>\u00a0makes it easy to take Step\u00a03.7\u00a0Flash\u00a0from development into production. Available as optimized, containerized inference microservices, NIM packages the model with the performance tuning, standardized APIs, and deployment flexibility enterprises need. Download and\u00a0run it\u00a0on-premises, in the cloud, or across hybrid environments.\u00a0NIM provides\u00a0a\u00a0standard\u00a0OpenAI inference\u00a0for sending inference\u00a0requests to\u00a0the NIM server.\u00a0<\/p>\n<p>Download the NIM container from\u00a0the\u00a0<a href=\"https:\/\/catalog.ngc.nvidia.com\/orgs\/nim\/teams\/stepfun-ai\/containers\/step-3.7-flash\" target=\"_self\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"internal\">NVIDIA container registry<\/a>\u00a0(enterprise license\u00a0required).\u00a0<\/p>\n<p>Start a server with the OpenAI\u00a0client.\u00a0<\/p>\n<p>Send either text or image\u00a0input\u00a0to\u00a0the endpoint.\u00a0<\/p>\n<p>from openai import OpenAI <\/p>\n<p>client = OpenAI(<br \/>\n  base_url = &#8220;http:\/\/0.0.0.0:8000\/v1&#8243;,<br \/>\n  api_key=&#8221;no-key-required&#8221;<br \/>\n) <\/p>\n<p>completion = client.chat.completions.create(<br \/>\n  model=&#8221;stepfun\/step-3.7-flash&#8221;,<br \/>\n  messages=[{&#8220;role&#8221;:&#8221;user&#8221;,&#8221;content&#8221;:&#8221;Explain particle physics?&#8221;}]<br \/>\n  temperature=0.5,<br \/>\n  top_p=1,<br \/>\n  max_tokens=1024,<br \/>\n  stream=True<br \/>\n) <\/p>\n<p>for chunk in completion:<br \/>\n  if chunk.choices[0].delta.content is not None:<br \/>\n    print(chunk.choices[0].delta.content, end=&#8221;&#8221;)<\/p>\n<p>Day 0 fine-tuning with\u00a0NVIDIA\u00a0NeMo\u00a0Framework\u00a0<a href=\"#day_0_fine-tuning_with\u00a0nvidia\u00a0nemo\u00a0framework\u00a0\" aria-label=\"Scroll to Day 0 fine-tuning with\u00a0NVIDIA\u00a0NeMo\u00a0Framework\u00a0 section\" class=\"heading-anchor-link\"><\/a><\/p>\n<p>Step 3.7 Flash can be customized with domain-specific data using open libraries from\u00a0the\u00a0<a href=\"https:\/\/github.com\/NVIDIA-NeMo\/\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">NVIDIA NeMo framework<\/a>.\u00a0NVIDIA\u00a0<a href=\"https:\/\/github.com\/nvidia-nemo\/automodel\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">NeMo Automodel<\/a>\u00a0library\u00a0combines native\u00a0PyTorch\u00a0n-D parallelisms\u00a0with\u00a0optimized\u00a0performance and\u00a0supports Day 0 fine-tuning directly from Hugging Face model checkpoints without checkpoint conversion.\u00a0The\u00a0Automodel\u00a0<a href=\"https:\/\/github.com\/NVIDIA-NeMo\/Automodel\/blob\/main\/docs\/guides\/vlm\/step-3-7.md\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">fine-tuning recipe<\/a>\u00a0for Step 3.7\u00a0supports techniques such as supervised fine-tuning (SFT) and memory-efficient\u00a0LoRA\u00a0at 600 tokens\/sec on Hopper GPUs.\u00a0<\/p>\n<p>For\u00a0advanced\u00a0large-scale training, teams can also use the\u00a0<a href=\"https:\/\/github.com\/nvidia-nemo\/megatron-Bridge\/\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">NeMo Megatron-Bridge<\/a>\u00a0fine-tuning\u00a0<a href=\"https:\/\/github.com\/NVIDIA-NeMo\/Megatron-Bridge\/tree\/main\/examples\/models\/stepfun\/step37\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">recipe<\/a>,\u00a0which provides\u00a0additional\u00a0performance optimizations.\u00a0<\/p>\n<p>From data center deployments on NVIDIA Blackwell\u00a0to\u00a0deskside\u00a0with\u00a0<a href=\"https:\/\/www.nvidia.com\/en-us\/products\/workstations\/dgx-station\/\" target=\"_self\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"internal\">NVIDIA DGX Station<\/a>\u00a0to managed NIM microservices and Day 0 fine-tuning workflows, NVIDIA\u00a0provides\u00a0a range of options for integrating Step\u00a03.7\u00a0Flash\u00a0across\u00a0different stages\u00a0of development and deployment.\u00a0With\u00a0748 GB of coherent memory, DGX Station is ideal for\u00a0running Step 3.7 Flash with increased headroom for\u00a0the full\u00a0256k context length,\u00a0and faster\u00a0local\u00a0developer iteration.\u00a0<\/p>\n<p>NVIDIA\u00a0is an active contributor to the open-source ecosystem and has released several hundred\u00a0<a href=\"https:\/\/developer.nvidia.com\/open-source\" target=\"_self\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"internal\">projects under open source licenses<\/a>. NVIDIA is committed to open models such as Step 3.7\u00a0Flash\u00a0that promote AI transparency and enable users to share their AI safety and resilience work.\u00a0<\/p>\n<p>To get started, check out\u00a0<a href=\"https:\/\/huggingface.co\/stepfun-ai\/Step-3.7-Flash\" target=\"_blank\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"external\">Step\u00a03.7 Flash<\/a>\u00a0on Hugging Face,\u00a0test\u00a0it\u00a0with\u00a0your own\u00a0data on\u00a0build.nvidia.com, or\u00a0locally on DGX Station\u00a0using the\u00a0<a href=\"https:\/\/build.nvidia.com\/station\/vllm\" target=\"_self\" rel=\"noreferrer noopener follow nofollow\" data-wpel-link=\"internal\">vLLM\u00a0Playbook<\/a>.\u00a0<\/p>\n","protected":false},"excerpt":{"rendered":"AI\u00a0applications are\u00a0moving\u00a0beyond text generation to\u00a0multimodal\u00a0systems that can perceive, search, and reason\u00a0across images, documents,\u00a0video,\u00a0and language in real time\u2014turning\u00a0fragmented\u00a0information into&hellip;\n","protected":false},"author":2,"featured_media":465733,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-465732","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/465732","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=465732"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/465732\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/465733"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=465732"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=465732"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=465732"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}