{"id":542605,"date":"2026-03-17T15:59:08","date_gmt":"2026-03-17T15:59:08","guid":{"rendered":"https:\/\/www.newsbeep.com\/ca\/542605\/"},"modified":"2026-03-17T15:59:08","modified_gmt":"2026-03-17T15:59:08","slug":"netflix-found-a-faster-way-to-load-containers","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ca\/542605\/","title":{"rendered":"Netflix Found a Faster Way to Load Containers"},"content":{"rendered":"<p>The initial appeal with containers was hardware-agnosticism. What runs on your machine runs on production, as long as both ran on x86 CPUs. This interoperability is a big factor in scalability of Kubernetes. <\/p>\n<p>But when speed really matters, understanding what hardware to run can make a difference.<\/p>\n<p><a href=\"https:\/\/youtu.be\/Fojn5NFwaw8\" target=\"_blank\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/03\/Techstrong-Gang-Youtube-PodcastV2-770.png\" alt=\"Techstrong Gang Youtube\"\/><\/a><\/p>\n<p>When Netflix upgraded its container runtime, from Docker to the open source <a href=\"https:\/\/containerd.io\/\" rel=\"nofollow noopener\" target=\"_blank\">containerd<\/a>, it noticed some nodes were stalling out when operations scaled up.<\/p>\n<p>This was a major problem. When a Netflix user decides what to play on Netflix, it triggers hundreds of containers. So it was quite important to ruthlessly tack down even the tiniest performance bottleneck.<\/p>\n<p>The culprit turned out to be the way Netflix initialized containers. But the bug only showed itself in certain cases. With a bit of sleuthing, Netflix engineers found the delay was worse on some CPU architectures more than others, according to a blog post.\u00a0<\/p>\n<p>\u201cUnderstanding and optimizing both the software stack and the hardware it runs on is key to delivering seamless user experiences at Netflix scale,\u201d wrote Netflix Senior Software Engineer <a href=\"https:\/\/www.linkedin.com\/in\/andrew-halaney\/\" rel=\"nofollow noopener\" target=\"_blank\">Andrew Halaney<\/a> and Netflix Senior Performance Engineer <a href=\"https:\/\/www.linkedin.com\/in\/harshad-sane-56711a11\/\" rel=\"nofollow noopener\" target=\"_blank\">Harshad Sane<\/a>, in an engineering blog post,\u00a0 \u201c<a href=\"https:\/\/netflixtechblog.com\/mount-mayhem-at-netflix-scaling-containers-on-modern-cpus-f3b09b68beac\" rel=\"nofollow noopener\" target=\"_blank\">Mount Mayhem at Netflix: Scaling Containers on Modern CPUs.<\/a>\u201c<\/p>\n<p>Long Boot Times<\/p>\n<p>Like most Kubernetes operations, a Netflix application will scale until it maxes out a node, then it procures a new instance and continues scaling on that one.<\/p>\n<p>After about 100 containers, however, Netflix was finding that their servers <a href=\"https:\/\/github.com\/containerd\/containerd\/issues\/12048#issuecomment-3050444019\" rel=\"nofollow noopener\" target=\"_blank\">were starting to slow<\/a>.\u00a0 A health check that read the mount table (a list of everything mounted on a server) would take 30 seconds or longer to complete. This was problematic in that Linux has the CPU dedicate itself to checking if the health check was complete, delaying everything else in processing queue.<\/p>\n<p>Too Many UIDs<\/p>\n<p>Under the older Docker setup, containers typically shared the host\u2019s user ID (UID) pool. When Netflix migrated to containerd\u2014the runtime responsible for executing the Kubelet\u2019s management tasks\u2014it shifted to a more secure User Namespace model. In this setup, each container is isolated with its own unique UID range.<\/p>\n<p>However, this approach required the kernel to individually \u2018idmap\u2019 every layer of a container image to that specific range. For a multi-layered image, this meant a massive spike in kernel calls just to instantiate a single container.<\/p>\n<p>Containerd, when assembling a container\u2019s root filesystem, is very needy with its requests for kernel-level locks, dominates the CPU\u2019s time \u2014 especially for containers with more than 50 layers.<\/p>\n<p>\u201cIf a node is starting many containers at once, every CPU ends up busy trying to execute these mounts,\u201d the pair of Netflix engineers wrote. \u201cAny system trying to quickly set up many containers is prone to this, and this is a function of the number of layers in the container image.\u201d<\/p>\n<p>Spinning up 100 containers, for instance, would require 20,200 mounts, each requiring system calls to the kernel!<\/p>\n<p>Multi-Core Engineering<\/p>\n<p>Netflix is not alone in suffering from this issue.\u00a0 Meta engineers, working with thousands of containers for AI inferencing, have <a href=\"https:\/\/www.socallinuxexpo.org\/scale\/23x\/presentations\/containers-all-way-down-what-we-learned-running-containers-containers-meta\" rel=\"nofollow noopener\" target=\"_blank\">bemoaned the performance drops<\/a> that came along with managing containers with thousands of UIDs.<\/p>\n<p>In fact, this has been a common system design issue for multi-core processing. Too often, data structures become the bottleneck, argued one <a href=\"https:\/\/pdos.csail.mit.edu\/papers\/corey:osdi08.pdf\" rel=\"nofollow noopener\" target=\"_blank\">widely-cited paper<\/a> for the 2008 USENIX Operating System and Design conference. Those researchers<a href=\"https:\/\/www.usenix.org\/legacy\/events\/osdi10\/tech\/full_papers\/Boyd-Wickizer.pdf\" rel=\"nofollow noopener\" target=\"_blank\"> also found<\/a> that many commonly-used applications \u2014 including Exim, memcached, Apache, PostgreSQL, and MapReduce \u2014 also suffer from kernel-level bottlenecks in multi-core servers.<\/p>\n<p>Linux kernel developer Christian Brauner, creator of the next-generation <a href=\"https:\/\/www.phoronix.com\/news\/Linux-7.0-Dropping-Old-Mount\" rel=\"nofollow noopener\" target=\"_blank\">Virtual File System mounting<\/a>, has<a href=\"https:\/\/brauner.io\/2023\/02\/28\/mounting-into-mount-namespaces.html\" rel=\"nofollow noopener\" target=\"_blank\">\u00a0long argued<\/a> that the mount() system call is broken for modern containers, suggesting that file-descriptor-based mounting could be used instead.<\/p>\n<p>The CPU Bottleneck<\/p>\n<p>But it gets even weirder when the engineering team looked at the difference between the hardware running these containers!<\/p>\n<p>Predominately, these timeouts were happening with the Intel Xeon-based AWS <a href=\"https:\/\/aws.amazon.com\/ec2\/instance-types\/r5\/\" rel=\"nofollow noopener\" target=\"_blank\">r5.metal<\/a> instances (with the Intel Skylake\/Cascade Lake architecture with 96 virtual CPUs). They happened far less frequently on either the later 7th generation Intel <a href=\"https:\/\/aws.amazon.com\/ec2\/instance-types\/m7i\/\" rel=\"nofollow noopener\" target=\"_blank\">m7i.metal-24xl<\/a> or the AMD EPYC -based <a href=\"https:\/\/aws.amazon.com\/ec2\/instance-types\/m7a\/\" rel=\"nofollow noopener\" target=\"_blank\">m7a.24xlarge<\/a>, which were also used on the Netflix Kubernetes deployments.<\/p>\n<p>The engineering team developed a <a href=\"https:\/\/github.com\/Netflix\/global-lock-bench\" rel=\"nofollow noopener\" target=\"_blank\">microbenchmark<\/a> to compare the lock contention across different multi-core systems. They found that the older mesh architecture used in r5.metal chips was the bottleneck. The design struggled to synchronize the global mount lock across multiple cores, leading to massive cache-line contention.<\/p>\n<p>The other AWS instances use a distributed architecture, where multiple cores each have their own local last-level cache. Lock contention is more rare in these designs.<\/p>\n<p>\u201cCentralized cache management amplified cache contention while distributed cache design smoothly scaled under load,\u201d the engineers concluded.<\/p>\n<p>The Fix Goes Upstream<\/p>\n<p>The Netflix engineering team tackled the issue at the software level, namely by reducing the number of kernel system calls the containerd was making.<\/p>\n<p>They created a <a href=\"https:\/\/github.com\/containerd\/containerd\/pull\/12092\" rel=\"nofollow noopener\" target=\"_blank\">pull request<\/a> that changed how containerd did global lock usage. This was made possible by Linux kernel 6.3, released in April 2023, which introduced support for recursive binds in\u00a0mount\u2019s rbind option.<\/p>\n<p>Instead of requiring global locks to mount each layer, containerd simply performed <a href=\"https:\/\/github.com\/containerd\/containerd\/pull\/12092\" rel=\"nofollow noopener\" target=\"_blank\">one single recursive bind mount<\/a> (idmap) of the entire parent directory where all those layers reside.<\/p>\n<p>\u201cThis makes the number of mount operations go from O(n) to O(1) per container, where n is the number of layers in the image,\u201d the engineers wrote.<\/p>\n<p>The PR, which included the metrics from the microbenchmark, was merged into <a href=\"https:\/\/newreleases.io\/project\/github\/containerd\/containerd\/release\/v2.2.0\" rel=\"nofollow noopener\" target=\"_blank\">containerd version 2.2<\/a>, released in November.<\/p>\n<p>But Netflix also considered the hardware, opting to route workloads away from r5.metal and to the architectures that scaled better under these conditions.<\/p>\n<p>\u201cThis experience underscores the importance of holistic performance engineering,\u201d they concluded.<\/p>\n<p>\n\tRelated<\/p>\n","protected":false},"excerpt":{"rendered":"The initial appeal with containers was hardware-agnosticism. What runs on your machine runs on production, as long as&hellip;\n","protected":false},"author":2,"featured_media":542606,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[49,48,211769,27784,25379,1316,61],"class_list":["post-542605","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-ca","tag-canada","tag-containerd","tag-containers","tag-kubernetes","tag-netflix","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/542605","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/comments?post=542605"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/542605\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media\/542606"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media?parent=542605"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/categories?post=542605"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/tags?post=542605"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}