{"id":404620,"date":"2026-04-30T07:49:13","date_gmt":"2026-04-30T07:49:13","guid":{"rendered":"https:\/\/www.newsbeep.com\/nz\/404620\/"},"modified":"2026-04-30T07:49:13","modified_gmt":"2026-04-30T07:49:13","slug":"noc-coherency-challenges-balloon-with-ai-socs-and-chiplets","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/nz\/404620\/","title":{"rendered":"NoC Coherency Challenges Balloon With AI SoCs And Chiplets"},"content":{"rendered":"<p style=\"font-weight: 400;\">Key Takeaways<\/p>\n<p>Data movement, congestion, and energy efficiency are key determiners of whether compute is usable.<br \/>\nDifferent processors bring various coherency challenges. For example, a cache-coherent NoC for CPUs is expensive and harder to verify than an I\/O-coherent NoC for an accelerator.<br \/>\nDesigners need to balance top-down performance with bottom-up physical engineering to effectively manage multiple NoCs and sub-NoCs.<\/p>\n<p style=\"font-weight: 400;\">Moving, managing, and keeping track of data is becoming a much bigger challenge as the amount of data that needs to be processed, stored, and accessed by different processing elements continues to grow.<\/p>\n<p style=\"font-weight: 400;\">Complex SoCs and multi-die implementations, particularly those involving AI, may contain numerous networks-on-chip (NoCs) to manage and prioritize data movement. They can include coherent or non-coherent caches and I\/Os, or they can just handle a single physical portion of the system. But all of this needs to be planned much earlier in the design flow than in the past, and it needs to be monitored throughout a system\u2019s lifecycle.<\/p>\n<p style=\"font-weight: 400;\">\u201cFrom what we\u2019re hearing across AI SoC teams, training and inference haven\u2019t just increased data volumes \u2014 they\u2019ve exposed data movement itself as the dominant system constraint,\u201d said Andy Nightingale, vice president of product management and marketing at <a href=\"https:\/\/semiengineering.com\/entities\/arterisip\/\" rel=\"nofollow noopener\" target=\"_blank\">Arteris<\/a>. \u201cCompute capability is scaling far faster than Moore\u2019s Law, but data movement, congestion, and energy efficiency increasingly determine whether that compute is usable at all.\u201d<\/p>\n<p style=\"font-weight: 400;\">Data traffic and coherence are good starting points for determining the best NoC for a particular use case. \u201c[Chip architects should] ask who needs coherence and why, which agents generate bursts versus steady streams, where latency bounds really matter, and how much reuse or scaling is expected across derivatives or chiplets,\u201d Nightingale said. \u201cCPU clusters tend to require coherent NoCs because their programming model depends on them. NPUs are usually non-coherent because explicit data movement and local memory deliver better power and throughput.\u201d<img data-recalc-dims=\"1\" fetchpriority=\"high\" decoding=\"async\" class=\"alignnone size-full wp-image-24275854\" src=\"https:\/\/www.newsbeep.com\/nz\/wp-content\/uploads\/2026\/04\/Arteris-SoC-with-CPU-NPU-NoC.png\" alt=\"\" width=\"1290\" height=\"434\"  \/><\/p>\n<p style=\"font-weight: 400;\">Fig. 1: CPU vs. NPU networks on chip. Source: Arteris<\/p>\n<p style=\"font-weight: 400;\">Others concur that coherency is a good starting point. \u201cChoose a coherent NoC for shared-memory CPU clusters where consistency matters, and a non-coherent NoC for NPUs and accelerators where throughput matters more than strict coherency,\u201d said William Wang, CEO of <a href=\"https:\/\/semiengineering.com\/entities\/alpha-design-ai-chipagents\/\" rel=\"nofollow noopener\" target=\"_blank\">ChipAgents<\/a>.<\/p>\n<p style=\"font-weight: 400;\">NoCs come in multiple flavors. They can be fully cache-coherent, last-level cache coherent, I\/O coherent (also known as one-way coherent), or non-coherent.<\/p>\n<p style=\"font-weight: 400;\"><img loading=\"lazy\" data-recalc-dims=\"1\" decoding=\"async\" class=\"alignnone size-full wp-image-24275855\" src=\"https:\/\/www.newsbeep.com\/nz\/wp-content\/uploads\/2026\/04\/Arteris-Coherent-NoC.png\" alt=\"\" width=\"1328\" height=\"508\"  \/><\/p>\n<p style=\"font-weight: 400;\">Fig. 2: Examples of coherent NoC deployments. Source: Arteris<\/p>\n<p style=\"font-weight: 400;\">Coherent networks tend to be more expensive and power-hungry than non-coherent networks. \u201cA very common paradigm is to take the big, powerful CPU that has these caches, and connect them up in a coherent network while trying to keep the coherent part as small as possible,\u201d said Kent Orthner, principal solutions architect at <a href=\"https:\/\/semiengineering.com\/entities\/baya-systems\/\" rel=\"nofollow noopener\" target=\"_blank\">Baya Systems<\/a>. \u201cYou really want to keep that only between the memory and the CPUs, and maybe AI accelerators. The rest of the system will typically use a simpler protocol that just reads and writes. The simple protocol doesn\u2019t have any notion of who last touched the data or who\u2019s responsible for it. They just go to an endpoint like a memory or PCIe control, or something like that, and test the data that they need. One common breakdown when people talk about multiple NoCs on an SoC is the coherent versus non-coherent domain.\u201d<\/p>\n<p style=\"font-weight: 400;\">In the past, large chipmakers often developed their own relatively simple NoCs. But as the complexity of data movement has grown, the industry has migrated toward commercially developed NoC IP. \u201cThey each have their approaches to the question, \u2018How can I provide you configurable blocks based on the input requirements?\u2019\u201d said Frank Schirrmeister, executive director for strategic programs and system solutions at <a href=\"https:\/\/semiengineering.com\/entities\/synopsys-inc\/\" rel=\"nofollow noopener\" target=\"_blank\">Synopsys<\/a>. \u201cYou say, \u2018Here are all my sources. I have 30 sources. Some are coherent, non-coherent, cache-coherent, and some are I\/O-coherent only.\u2019 Then you get into the environment where you configure the NoC, push a button, and they have building blocks, being able to instantiate these caches, TLBs (translation lookaside buffers), and the like, to then build a NoC. When it comes to the non-coherent bits, the focus is on making sure they can be implemented. Looking at the layout, what is the timing? Where are these things? It is quite messy on those pesky multi-billion-gate chips to determine where to put things and how to carry data back and forth. The challenge is significant.\u201d<\/p>\n<p style=\"font-weight: 400;\"><img loading=\"lazy\" data-recalc-dims=\"1\" decoding=\"async\" class=\"alignnone size-full wp-image-24275856\" src=\"https:\/\/www.newsbeep.com\/nz\/wp-content\/uploads\/2026\/04\/Arteris-non-coherent-NoC.png\" alt=\"\" width=\"1146\" height=\"556\"  \/><\/p>\n<p style=\"font-weight: 400;\">Fig. 3: Examples of non-coherent NoC deployments. Source: Arteris<\/p>\n<p style=\"font-weight: 400;\">Cache coherency is more challenging than I\/O coherency. \u201cThe space is much bigger for I\/O core coherency,\u201d said Schirrmeister. \u201cYou typically have the peripheral, such as a GPU or network interface, and you read the CPU cache data, but not vice versa. But in cache coherency, you have multiple CPU cores that need to see a consistent view of the memory across all of them, which is much more complicated because there are more options.\u201d<\/p>\n<p style=\"font-weight: 400;\">Caches are used to temporarily store data close to the processor. \u201cThe processor says, \u2018I need this, I need that, and I need this other thing,\u2019\u201d said Orthner. \u201cIt doesn\u2019t have to go all the way to external memory, which can take 100 nanoseconds. It can get it really quickly. But with the big systems that we tend to look at, you could have many, many different processors. And as they share memory, if you have one processor saying, \u2018Look, I need this little bit of information,\u2019 and it doesn\u2019t know that another processor already has a copy of it locally in its cache, then it can end up getting the wrong value. So cache coherency is all about making sure that each processor\u2019s view of the information is the same cache.\u201d<\/p>\n<p style=\"font-weight: 400;\">In fact, each type of compute processor has its own private cache of data. \u201cThey have to keep track of who has what piece of data and share it efficiently, and that\u2019s what the cache coherency protocol was all about,\u201d said Orthner. \u201cIf you have an SoC with 500 processors, it knows what those 500 processors are and how to talk to them, because all of that was decided before you manufactured anything.\u201d<\/p>\n<p>Chiplets, multi-die, and 3D<br \/>Multi-die assemblies require additional management of coherent and non-coherent NoCs.<br \/>\u201cOur physical AI chiplet platform is typically a minimum of three chiplets,\u201d said Mick Posner, senior product marketing group director for chiplets &amp; IP solutions at <a href=\"https:\/\/semiengineering.com\/entities\/cadence-design-systems\/\" rel=\"nofollow noopener\" target=\"_blank\">Cadence<\/a>. \u201cIn the center, you have the system chiplet. On one side you have a CPU chiplet. On the other side you have an AI accelerator. That\u2019s the base three chiplets of any physical AI chiplet platform, and that center system chiplet must have both a coherent interface and a non-coherent interface because it will talk to a CPU. So it must have coherency between the system and the CPU, because it\u2019s managing the memory. It also needs a coherent interface to it. But the link to the AI accelerator is only I\/O coherent. It doesn\u2019t require cache coherency, because the accelerator is like an extension. You\u2019re just sending something to it. It typically can have its own memory, maybe it\u2019s sharing memory, but it doesn\u2019t need cache coherency.\u201d<\/p>\n<p style=\"font-weight: 400;\">Multiple NoCs are needed to achieve this. \u201cA NoC that is designed for coherency usually has a lot more overhead than a non-coherent one,\u201d said Posner. \u201cThey probably have a link between them, but that link is non-coherent by default because there\u2019s no coherency on one side.\u201d<\/p>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-24275857\" src=\"https:\/\/www.newsbeep.com\/nz\/wp-content\/uploads\/2026\/04\/Arteris-coherent-and-non-coherent-NoC-1.png\" alt=\"\" width=\"1378\" height=\"396\"  \/><\/p>\n<p style=\"font-weight: 400;\">Fig. 4: SoCs can contain a mixture of non-coherent and coherent NoCs. <a href=\"https:\/\/www.arteris.com\/blog\/the-soc-interconnect-fabric-a-brief-history\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>: Arteris<\/p>\n<p>Balancing coherency, programmability, and integration<br \/>Orchestrating data movement in multi-die assemblies puts added demands on programmability and system discovery.<\/p>\n<p style=\"font-weight: 400;\">NoCs can help chiplets find each other in the package. \u201cIn a chiplet approach you might have all the same chip, on the same die, the same piece of silicon, but you might have it by itself in a package,\u201d said Baya\u2019s Orthner. \u201cYou might have it with an array of four other chiplets in a package. You might stack it vertically in a different package. So now you have this whole interesting problem of, when you first power up, how do the chiplets discover each other, and how do they learn who I am in the context of this package, and where is everybody? As a result, you end up needing a much higher degree of programmability.\u201d<\/p>\n<p style=\"font-weight: 400;\">At boot-up, systems need a management agent. \u201cIt comes in and says, \u2018If you\u2019re looking for this subset of processes, you\u2019re going to have to send your traffic to the north side of the die, because that\u2019s where they\u2019re located \u2014 unless you\u2019re the chiplet that\u2019s on the north, because then you\u2019re going to have to send your stuff south,\u2019\u201d Orthner said. \u201cYou end up waking up, communicating, discovering, and then reconfiguring the routing in the network so that different chiplets can recognize the same destination, but via different routes.\u201d<\/p>\n<p style=\"font-weight: 400;\">Stacked-die configurations add more networking challenges. \u201cWith the different vendors in a die, how are you going to manage the different types of parts and determine how well they are integrated into the system?\u201d said Hee Soo Lee, high-speed digital design segment lead at <a href=\"https:\/\/semiengineering.com\/entities\/keysight-technologies\/\" rel=\"nofollow noopener\" target=\"_blank\">Keysight EDA<\/a>. \u201cThese challenges are making networks and I\/Os complicated. With stacked die configurations, managing the thermal and mechanical issues will be more significant problems than electrical issues. All of these systems are driven by the market, especially the ever-increasing demand for data in data center AI workloads.\u201d<\/p>\n<p>Optimizing for PPA<br \/>Whether coherent or non-coherent, multiple NoCs can be considered from a top-down approach to chip design.<\/p>\n<p style=\"font-weight: 400;\">\u201cLet\u2019s say you have two at the main level, with the top ones connecting all the subsystems,\u201d said Cadence\u2019s Posner. \u201cUnderneath those, there are typically localized networks on chips. There could even be multiple \u2014 one for I\/O peripherals that is low-bandwidth, low-performance, and a separate one for high-bandwidth peripherals. You want to tailor your NoC to whatever you\u2019re connecting it to and configure it for power, performance, and area (PPA) based on the network that it\u2019s connecting to. You need to understand that hierarchy because it all still needs to be controlled.\u201d<\/p>\n<p style=\"font-weight: 400;\">AI-driven EDA technology is helping to sort through PPA tradeoffs. \u201cTools are evolving toward a fully autonomous, multi-agent workflow operating system that reasons across spec-to-silicon with real-time design-quality feedback, PPA optimization, and cross-domain co-design,\u201d noted ChipAgents\u2019 Wang.<\/p>\n<p style=\"font-weight: 400;\">In large, powerful SoCs, designers will often create several configurations. \u201cTypically, they want the configuration network to be separate from their coherent and primary non-coherent data paths, and configuration is pretty low throughput,\u201d said Baya\u2019s Orthner. \u201cIt\u2019s your control plane. It wants to read registers, check up on the performance of the system, maybe manage some power control stuff where you shut down parts of the device and turn it on again. That\u2019s also typically implemented as a distinct network. It\u2019s the same protocols as you would see in your data flow network, but on a much smaller scale. You want to be able to access it without affecting your primary data traffic, stealing some of the network bandwidth, for instance.\u201d<\/p>\n<p style=\"font-weight: 400;\">Geographical or topological differences also drive the need for distinct NoCs. \u201cYou might have, in the west part of your chip, an array of compute cores,\u201d said Orthner. \u201cIn the east part of your chip, you might have a different array of compute cores, and for your primary data paths, they do not need to talk to each other, so you can implement that as a west network, an east network, and a third network that brings them all together.\u201d<\/p>\n<p style=\"font-weight: 400;\">Memory requirements are forcing new network design decisions to enable high-performance systems. \u201cEven though they\u2019re geographically the same location, people will take half the memory space and call it the \u2018even\u2019 network, with the other half the \u2018odd\u2019 network, just so they can get twice as much data flowing around the system without stepping on its toes,\u201d Orthner explained. \u201cThere are many ways to have different logical networks. When you\u2019re designing a chip, you think, \u2018Here\u2019s the NoC, and here\u2019s the area that it represents.\u2019 Then the physical design engineers want to place all of the pieces of that inside that rectangle and get it to meet timing. If you\u2019ve used different tools, or done different projects with the same tool to do all of your different NoCs, you\u2019d end up with a lot of logic with different locations in the hierarchy that are superimposed on each other, which makes it very difficult for the physical designer.\u201d<\/p>\n<p style=\"font-weight: 400;\">These distinct network approaches underscore the complexity of modern chip design, especially as physical and logical requirements diverge. To address these challenges and maintain efficiency, designers are increasingly looking for unified solutions that seamlessly integrate multiple NoCs within a single project.<\/p>\n<p style=\"font-weight: 400;\">Unified NoC software refers to a way of managing the overall project of multiple NoCs, rather than a confluence of all the data in a single NoC. \u201cWhen we talk about a unified network, what we\u2019re saying is you can have one design, one top-level project, one idea, for the routers and the logic and the different tracks and the positioning and everything else,\u201d Orthner explained. \u201cAnd within the design of the big picture network, we\u2019re keeping all of the smaller networks distinct. If you look at coherent versus non-coherent, you can say, \u2018I know enough about my traffic flow to say I want to allow my non-coherent traffic and my coherent traffic to be logically separate, but still use thin wires,\u2019 which is a really strong capability nowadays when you\u2019re trying to optimize for area and cost and everything else. You can have one run of the tool that embraces all of the different networks that are running in parallel, and which allows them to share resources under the system designer\u2019s control.\u201d<\/p>\n<p style=\"font-weight: 400;\">As chip complexity continues to rise, it becomes increasingly important to consider how these unified NoC strategies translate to larger system architectures. Bridging the gap between on-chip networks and broader infrastructure, designers also must address how these concepts scale up to the data center level.<\/p>\n<p>Data center hierarchy<br \/>Considering NoCs from the perspective of a data center, several fundamental questions arise. How do the racks communicate with one another? How do the cards within a rack communicate? What methods allow chips on a card to interact? And finally, how is communication achieved between chiplets?<\/p>\n<p style=\"font-weight: 400;\">Answering these questions provides the top-down view of the data center, which helps to establish hierarchy between the numerous NoCs in use. \u201cThe top level is really the data center level, but you start at the data because the whole data center acts as your computer at this point,\u201d said Saurabh Gayen, chief solutions architect at Baya Systems. \u201cYou need to make sure that you understand how the racks are organized, how data is flowing across the scale-out domain, then, going into the scale-up domain, and further hierarchically going down through your particular package level, into the chiplet level, etc. You have to create that top-down view because that defines the hierarchical design and how to group things.\u201d<\/p>\n<p style=\"font-weight: 400;\">While it\u2019s easier to build many small things that are loosely connected, that\u2019s not the best approach. \u201cYou need much denser, tightly packed, smaller-level, flatter, higher-performance things that are tightly coupled and hierarchically organized together,\u201d said Gayen. \u201cYou cannot cheat by having tons of networks. The old data centers could do that. We would see a lot more scalar design, and the AI models were stalling. Now, as an industry, we bit the bullet and said, \u2018Yep, we\u2019re going to go all in on this to make our hierarchies flatter, higher performance, and denser.\u2019\u201d<\/p>\n<p style=\"font-weight: 400;\">These same hierarchy-level concepts apply as designers go down into packages and chiplets. \u201cYou take a top-down view,\u201d said Gayen. \u201cWe don\u2019t want to think about bottom-up, \u2018Here\u2019s a NoC and here\u2019s a NoC and here\u2019s a NoC and here\u2019s a NoC, and how do we switch them up together?\u2019 The best way is to look top-down at the overall system, and how to break it up into these smaller things. There are also differences between NoCs on a particular chiplet versus how they communicate with each other, die-to-die. The performance is top down, but the reality of engineering is bottom up, so you have to balance the two things hierarchically.\u201d<\/p>\n<p style=\"font-weight: 400;\">Correct data management comes down to the sub-NoC hierarchy. \u201cIn terms of the hierarchy within the design of a piece of silicon, how do you construct the sub NoCs, and the NoCs talking to each other, and the boundaries?\u201d Orthner said. \u201cIt is important to get it right, all the way down to physical design, when you\u2019re placing the transistors on the silicon, and take advantage of the hierarchy that you thought through, to minimize the effort that goes into it. Hierarchy is something that, when designing the framework for NoC design, we take as a first-class concern.\u201d<\/p>\n<p style=\"font-weight: 400;\">For performance analysis, it\u2019s especially important to look at the top level. \u201cInstead of thinking about how fast the memory interface is, you think about it in terms of the big picture data flows.\u201d said Orthner. \u201cWho needs to talk to who and why? Where\u2019s the data going to be moving? Almost always you have multiple data sources running at the same time, so, how are those going to affect each other and impact each other?\u201d<\/p>\n<p>Conclusion<br \/>In the past, NoCs often were an afterthought in chip design, addressed once the physical layout was nearly complete. Today, designers must integrate NoCs into the early stages of development, ensuring that communication infrastructure is optimized alongside processing elements. This approach improves performance, power efficiency, and scalability, allowing complex systems to achieve better overall results.<\/p>\n<p style=\"font-weight: 400;\">\u201cWhat\u2019s changing in practice is a reframing of priorities,\u201d said Arteris\u2019 Nightingale. \u201cData movement is now a first-class design axis alongside compute and memory, especially as systems scale from monolithic dies to chiplets and distributed architectures. Leading teams are investing earlier in architectures that provide visibility, quality-of-service guarantees, and long-term scalability, without treating interconnect as a late-stage optimization.\u201d<\/p>\n<p style=\"font-weight: 400;\">To keep pace with evolving system architectures, the industry must prioritize the development of robust, standardized security protocols and invest in scalable interconnect solutions from the earliest stages of design. Collaboration across the supply chain is essential to ensure trust, interoperability, and resilience against emerging threats, especially as chiplet and multi-die integration becomes more prevalent. Moving forward, ongoing research, cross-vendor partnerships, and the adoption of adaptive network topologies will be critical to meeting the demands of secure, high-performance data movement in next-generation systems.<\/p>\n<p>Related Article\u00a0<br \/><a href=\"https:\/\/semiengineering.com\/data-boom-puts-pressure-on-nocs-fabrics\/\" rel=\"nofollow noopener\" target=\"_blank\">Data Boom Puts Pressure On NoCs, Fabrics<\/a><br \/>New adaptive, mesh NoC topologies are enabling chip designers to optimize data movement in complex SoCs and multi-die systems.<\/p>\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"Key Takeaways Data movement, congestion, and energy efficiency are key determiners of whether compute is usable. Different processors&hellip;\n","protected":false},"author":2,"featured_media":404621,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[150854,166082,60292,208435,93763,166083,208436,208437,185151,185152,111,139,208438,69,93762,145],"class_list":["post-404620","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-arteris","tag-baya-systems","tag-cache","tag-cache-coherency","tag-cadence","tag-chipagents","tag-coherency","tag-hierarchy","tag-keysight-eda","tag-network-on-chip","tag-new-zealand","tag-newzealand","tag-noc","tag-nz","tag-synopsys","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts\/404620","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/comments?post=404620"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts\/404620\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/media\/404621"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/media?parent=404620"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/categories?post=404620"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/tags?post=404620"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}