{"id":415954,"date":"2026-04-24T23:00:13","date_gmt":"2026-04-24T23:00:13","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/415954\/"},"modified":"2026-04-24T23:00:13","modified_gmt":"2026-04-24T23:00:13","slug":"running-ai-on-vmware-workstation-virtualization-review","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/415954\/","title":{"rendered":"Running AI on VMware Workstation &#8212; Virtualization Review"},"content":{"rendered":"\n<p id=\"ph_pcontent1_0_KickerText\" class=\"kicker\"><a href=\"https:\/\/virtualizationreview.com\/articles\/list\/features.aspx\" rel=\"nofollow noopener\" target=\"_blank\">In-Depth<\/a><\/p>\n<p>        Running AI on VMware Workstation<\/p>\n<p>In a previous <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/16\/ai-on-a-raspberry-pi-part-3-testing-different-llms.aspx\" target=\"_blank\" rel=\"nofollow noopener\">article<\/a>, I ran LLMs on a Raspberry Pi 5, and although I was able to get it working in less than 5 minutes, it took around 17 minutes to run even the most basic task, thereby making it useless for real-time work. I hoped the limitation of running AI locally was due to the Pi&#8217;s underlying hardware, not to running AI locally. To test this, I will run LLMs on my laptop in a virtual machine to test their performance.<\/p>\n<p>Laptop Specifications<br \/>The laptop that I will be running it on is an older HP Firefly with an Intel Core i7-8665U. The CPU is an 8th-generation Whiskey Lake mobile processor launched in 2019, designed for laptops and other mobile devices. It has 4 cores, 8 threads, a 1.9 GHz base frequency, and 4.8 GHz in turbo mode. It has 8 MB Intel Smart Cache and integrated UHD Graphics 620, with a 15W TDP. It was designed to be a power-efficient CPU for business, multitasking, and office productivity. The system has 16 GB of RAM and a 477 GB SSD drive.<\/p>\n<p>Testing Framework<br \/>For these tests, I will be using VMware Workstation Pro 25H2, a totally free Type 2 hypervisor that runs on Windows or Linux. You can read more about <a href=\"https:\/\/virtualizationreview.com\/articles\/2024\/05\/29\/install-free-hypervisor.aspx\" target=\"_blank\" rel=\"nofollow noopener\">why I like it so much<\/a>.<\/p>\n<p>The VM I created to run my tests runs Ubuntu 24.04. I gave the VM 3 CPU cores and 12GB of RAM.\n<\/p>\n<p>As with my Pi testing, I will be testing the same small LLMs: gemma2:2b, tinyllama, deepseek-r1:1.5b, and qwen2.5:3B.\n<\/p>\n<p>I will run the same three tests on all the LLMs above using the same script that I used for my Pi testing.<\/p>\n<p>You can read more about these models and my testing script in my previous article.<\/p>\n<p>Testing the LLMs<br \/>I powered down all other applications on the Windows laptop while the tests were running.<\/p>\n<p>While the tests were running, I monitored the laptop using Task Manager and the VM using htop.<\/p>\n<p>The tests ran without a hitch. On the Windows machine, Task Manager showed that three cores were at 100% and that all 12GB of RAM was in use. On the Linux VM, htop showed that all three assigned cores were in use during the tests.<\/p>\n<p>Below is an example of the summary that it produced.<\/p>\n<p> <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_87670819.ashx\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" alt=\"Picture 1\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_87670819_s.ashx.jpeg\" width=\"300\" height=\"134\"\/> <\/a><br \/>\n [Click on image for larger view.]<\/p>\n<p>Test Results and Analysis<br \/>To get an accurate result, I ran my testing script 4 times. The test results were consistent across runs.<\/p>\n<p>I parsed the results for the Model, Prompt, total duration, prompt eval count, and eval count in the output and put them in a spreadsheet for further analysis.<\/p>\n<p> <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_637ff976.ashx\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" alt=\"Shape1\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_637ff976_s.ashx.jpeg\" width=\"300\" height=\"116\"\/> <\/a><br \/>\n [Click on image for larger view.]<\/p>\n<p>Digging into the data, I found that these results revealed performance patterns across LLMs. This highlights the importance of matching the model to the workload type.\n<\/p>\n<p>I found that TinyLlama consistently delivered the fastest response times. This would make it my choice for real-time inference, edge deployments, and latency-sensitive applications, particularly for factual queries and lightweight code generation.<\/p>\n<p>Qwen2.5:3B provided a balance of speed and output quality. It performed reliably across all tasks with moderate token usage and predictable runtimes. It would be well-suited for general-purpose workloads and RAG pipelines, which I hope to work with in the future.<\/p>\n<p>Gemma2:2B showed mixed behavior, excelling at simple factual queries but slowing significantly on structured tasks such as my HTML generation and tabular data creation prompts.<\/p>\n<p>I found it interesting that DeepSeek-R1:1.5B exhibited extreme variability. It was relatively fast on simple code generation but slowed down on reasoning tasks, with token counts and runtimes increasing by an order of magnitude. This reflected its chain-of-thought reasoning architecture, which prioritized analytical depth over speed.\n<\/p>\n<p>Overall, these benchmarks demonstrated to me that no single model is optimal for all workloads. Significant performance and cost efficiencies can be achieved through intelligent model choice: using lightweight models for routine tasks and reserving reasoning-focused models for complex analytical tasks, running them on hardware that supports them.<\/p>\n<p>Charting the Data<br \/>To help me visualize my data, I aggregated it and created a few charts.<\/p>\n<p>Average Token Output by Model<\/p>\n<p>This chart compares the average number of tokens generated by each LLM.<\/p>\n<p>I found that DeepSeek-R1 produced far more tokens than the other LLMs, confirming its heavy chain-of-thought reasoning style. The smaller LLMs were far more concise, improving speed and efficiency.<\/p>\n<p> <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_4c5cd45a.ashx\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" alt=\"Picture 22\" src=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_4c5cd45a_s.ashx\" width=\"300\" height=\"226\"\/> <\/a><br \/>\n [Click on image for larger view.]<\/p>\n<p>Compute Efficiency (Seconds per Token)<\/p>\n<p>This was a very revealing chart as it shows how much time each model spends per generated token.<\/p>\n<p>Looking at this data, I found that TinyLlama was the most efficient model. Qwen2.5:3B provided a strong balance of speed and output quality, while DeepSeek-R1 was extremely inefficient for throughput, as expected given its focus on reasoning.<\/p>\n<p> <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_7143ad31.ashx\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" alt=\"Picture 23\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_7143ad31_s.ashx.jpeg\" width=\"300\" height=\"226\"\/> <\/a><br \/>\n [Click on image for larger view.]<\/p>\n<p>VMware Workstation vs. Pi<br \/>I then pulled up the results from my Pi testing and compared them with those from my x64 VM. I found that my Pi was significantly slower than x86 systems.<\/p>\n<p>Among the models tested, TinyLlama demonstrated the most efficient scaling when run on Raspberry Pi devices. In contrast, DeepSeek-R1 proved unsuitable for reasoning tasks on ARM platforms due to the time it took to complete them.<\/p>\n<p> <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_3df578f5.ashx\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" alt=\"Picture 51\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_3df578f5_s.ashx.jpeg\" width=\"300\" height=\"227\"\/> <\/a><br \/>\n [Click on image for larger view.]<\/p>\n<p> <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_84bfcd87.ashx\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" alt=\"Picture 58\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_84bfcd87_s.ashx.jpeg\" width=\"300\" height=\"227\"\/> <\/a><br \/>\n [Click on image for larger view.]<\/p>\n<p> <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_3df578f5.ashx\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" alt=\"Picture 65\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_3df578f5_s.ashx.jpeg\" width=\"300\" height=\"227\"\/> <\/a><br \/>\n [Click on image for larger view.]<\/p>\n<p>The performance gap between the HP Firefly (Intel i7-8665U) and Raspberry Pi 5 (ARM BCM2712) is driven by a combination of architectural, microarchitectural, software, and memory-system factors.<\/p>\n<p>To get an idea of why there is such a wide disparity in performance between the systems, I went to <a href=\"https:\/\/www.cpu-monkey.com\" target=\"_blank\" rel=\"nofollow noopener\">cpu-monkey.com<\/a> and compared the CPUs in each system.<\/p>\n<p> <a href=\"https:\/\/virtualizationreview.com\/articles\/2026\/04\/24\/~\/media\/ECG\/virtualizationreview\/Images\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_3473644e.ashx\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" alt=\"Shape2\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/04\/RTP_VR_Running_AI_VMW_WS_P4_html_3473644e_s.ashx.jpeg\" width=\"300\" height=\"106\"\/> <\/a><br \/>\n [Click on image for larger view.]<\/p>\n<p>The performance gap between the HP Firefly laptop with its Intel Core i7-8665U and the Raspberry Pi 5 with its ARM BCM2712 processor is driven by differences in CPU architecture, vector processing capabilities, memory hierarchies, and software optimization.\n<\/p>\n<p>The Intel processor benefits from a wide, out-of-order execution pipeline, large caches, aggressive branch prediction, and advanced SIMD vector instructions (AVX2), all of which are heavily leveraged during LLM inference. These capabilities allow modern Intel CPUs to process matrix operations and token generation far more efficiently than the<br \/>\nPi&#8217;s ARM cores, which rely on smaller caches, lower memory bandwidth, and more limited vector units. As a result, even my smaller LLMs ran dramatically faster on my x64 laptop, with speed differences ranging from roughly 10\u00d7 to more than 90\u00d7!<\/p>\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"In-Depth Running AI on VMware Workstation In a previous article, I ran LLMs on a Raspberry Pi 5,&hellip;\n","protected":false},"author":2,"featured_media":415955,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[61,60,182431,80],"class_list":["post-415954","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-ie","tag-ireland","tag-raspberry-pi-ai","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/415954","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=415954"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/415954\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/415955"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=415954"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=415954"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=415954"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}