{"id":589761,"date":"2026-08-11T08:52:11","date_gmt":"2026-08-11T08:52:11","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/589761\/"},"modified":"2026-08-11T08:52:11","modified_gmt":"2026-08-11T08:52:11","slug":"hbm-becomes-testbed-for-3d-assembly-yield","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/589761\/","title":{"rendered":"HBM Becomes Testbed For 3D Assembly Yield"},"content":{"rendered":"<p>Key Takeaways:<\/p>\n<p>DFT is increasingly critical for detecting defects in memory cells, TSVs, microbumps, and high-speed interfaces, while power integrity and thermal effects add further test challenges.<br \/>\nMarginal or aging interconnects are caught by a combination of BiST, embedded monitors, boundary scan, at-speed testing, redundancy\/repair, and in-system monitoring, a complex hierarchy of testing.<br \/>\nHBM is a reliability bottleneck, requiring greater visibility into individual bonds and TSVs, and more sophisticated SI\/PI and repair mechanisms for latent failures in complex 2.5D\/3D AI systems.<\/p>\n<p>The high cost of field failures in data centers is driving big changes in design-for-test, enabling chipmakers to identify impending failures and root out interconnect and die-to-die weaknesses. This is especially important for high-bandwidth memory (HBM), which is acting as the frontier architecture for proving out 3D manufacturing and testing strategies.<\/p>\n<p>HBM is well known for the tremendous benefits it delivers to data-intensive tasks such as AI, high-resolution graphics processing, and other applications that need massive data processing and high-speed data transfer. Its tightly coupled, vertically stacked DRAM creates ultrawide communication channels (1024, 2048 bits) that outperform all other memory types. But testing HBM faces tremendous challenges due to its complex structure, which becomes considerably harder with each device generation.<\/p>\n<p>\u201cNow we are preparing for HBM 5 technology, which will increase that stack height to 24, which means even more data going back and forth. A lot of ones and zeros in very close proximity with no growth in footprint\/area,\u201d said Faisal Goriawalla, director of product management at <a href=\"https:\/\/semiengineering.com\/entities\/synopsys-inc\/\" rel=\"nofollow noopener\" target=\"_blank\">Synopsys<\/a>. \u201cSo as the pitch between these signals is squeezed, you have electrical interference challenges such as the victim aggressor situation, where one net toggling could cause the other nets in its vicinity to also flip, which was not intended to happen. These tight pitches make the DFT aspects of testing, such as diagnosis, much more complicated in multi-die technology.\u201d<\/p>\n<p>This is especially evident with microbumps and bonded interconnects, which often cannot be accessed directly. \u201cBy virtue of these interconnects being inside a multi-die package, the testing that can be performed is limited,\u201d said Noam Brousard, vice president of solutions engineering at <a href=\"https:\/\/semiengineering.com\/entities\/proteantecs\/\" rel=\"nofollow noopener\" target=\"_blank\">proteanTecs<\/a>. \u201cSo we use minuscule monitors on the end of the HBM that measure each signal and determine how close it is to failure, providing an eye diagram per lane with very high accuracy.\u201d<\/p>\n<p>Such preventive techniques are becoming increasingly attractive to data centers running large language models, whose interruption leads to significant losses.<\/p>\n<p>In HBM3, 12 or more DRAM chiplets are stacked atop a silicon interposer, interconnected by finely spaced through-silicon vias (TSVs) and connected by microbumps. As much effort goes into testing the TSVs, microbumps, and die-die interfaces as for the memory cells themselves. For this reason, DFT architectures have become critical for manufacturers of HBM modules, including SK hynix, Samsung, and Micron.<\/p>\n<p>The magnitude of this challenge cannot be overstated. \u201cHBM testing can be a major bottleneck due to the complexity of test program development,\u201d said Quoc Phan, technology enablement manager of 3DIC DFT and yield at <a href=\"https:\/\/semiengineering.com\/entities\/mentor-a-siemens-business\/\" rel=\"nofollow noopener\" target=\"_blank\">Siemens EDA<\/a>. \u201cCreating the specialized fault models and test algorithms necessary for detecting defects in TSVs, microbumps, and inter-die interfaces is an intricate process requiring deep expertise. Integrating these advanced tests seamlessly with the host processor\u2019s DFT architecture and ensuring reliable communication between them adds substantial layers of complexity to both test program development and the subsequent debugging phases.\u201d<\/p>\n<p>With the upcoming transition to hybrid bonding for advanced HBM4 processes, much more of the focus will be on ensuring the quality of each interconnect bond. \u201cAs we\u2019re stacking these devices, the coplanarity, the warpage, the bonding processes, and everything that we\u2019re doing to make these individual bonds from a C4 bump to a copper- to-copper bump or a die-to-die connection \u2014 each one of these bonds is critical, said Jack Lewis, CTO of <a href=\"https:\/\/semiengineering.com\/entities\/modus-test\/\" rel=\"nofollow noopener\" target=\"_blank\">Modus Test<\/a>. \u201cThe amount of precise information we can get about these individual bonds and the performance of the bonding processes will help our customers ramp and hit yield entitlement quickly.\u201d<\/p>\n<p>DFT\u2019s mission is to enable a consistent way of testing chips post-manufacturing. But with an HBM stack, it also needs to detect defects in the memory cells and periphery, TSVs, microbumps, die-die interfaces, base die circuitry, package-level interconnects, and high-speed I\/O interfaces to the accelerator such as the UCI Express.<\/p>\n<p>Though the strategy by any one chipmaker is proprietary, relationships among companies are shifting. \u201cDFT is developed as proprietary IP by each device manufacturer, so it is difficult to comment on specific implementations,\u201d said Jin Yokoyama, senior director of memory product marketing at <a href=\"https:\/\/semiengineering.com\/entities\/advantest-corporation\/\" rel=\"nofollow noopener\" target=\"_blank\">Advantest<\/a>. \u201cHowever, looking forward, the boundary of responsibilities between SoC end users and memory vendors may become increasingly blurred and complex, particularly with respect to how DFT is defined and partitioned.\u201d<\/p>\n<p>Part of the reason for this change is the adoption of custom HBM for specific applications.<\/p>\n<p>Reconfiguring for custom HBM<br \/>The transition from HBM3 to HBM4 includes the option of a custom logic base die that replaces the DRAM-built controller die of previous generations. Custom HBM is attractive because it allows the AI accelerator or GPU designer to optimize the memory stack for its specific workload. This is particularly important for AI training and inference, where performance is more often bandwidth-limited rather than compute-limited. The downside is that the longer development and qualification cycle is likely to limit custom HBM to the highest-volume applications.<\/p>\n<p><img data-recalc-dims=\"1\" fetchpriority=\"high\" decoding=\"async\" class=\"alignnone size-full wp-image-24279915\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/08\/Synopsys_Speeding-Down-Memory-Lane-Custom-HBM-fig2.webp.png\" alt=\"\" width=\"725\" height=\"374\"  \/><br \/>Fig. 1: The HBM DRAM stack. Source: Synopsys<\/p>\n<p>Importantly, custom HBM\u2019s modifications alter the testing landscape. \u201cCustom HBM gives SOC designers tremendous flexibility to configure the logic base die the way they want,\u201d said Goriawalla. So it means that if you are in a data center AI training environment where latency and throughput are very important factors, you can configure your HBM controller for those goals. But if you are in an AI inference type of application, where area and power are bigger concerns, then you can configure the HBM controller and logic differently. So this means we have to be thoughtful about DFT based on the use case scenario.\u201d<\/p>\n<p>The additional logic circuitry in custom HBM means there is more opportunity to use on-die monitors on this base die to help detect timing margin problems. \u201cWe look at the SoCs and HBMs as a system because there\u2019s interaction between the two,\u201d said proteanTecs\u2019 Brousard. \u201cFor instance, a lot of traffic coming in from the HBM might cause a current surge, which leads to a voltage droop. This will readjust, but that sudden voltage droop can cause failures, and a sensor with fast response time can capture that change. We want to have that kind of visibility, because the SoC is being affected by the HBM. So it is really the work of many monitors together that provide the rich dataset needed not only to identify problems, but to infer the reason behind it so that it can be properly mitigated.\u201d<\/p>\n<p>HBM faces limits of probe-ability<br \/>Interconnect bump pitch between DRAMs in HBM is now so tight (&lt;40\u00b5m, with 20 to 25\u00b5m microbumps), that it has become nearly impossible to probe the microbumps directly. Even if the bumps could be probed, the chances for damage are too high. So in many cases, larger sacrificial pads are used for probing, but in the long run it appears that engineers will be increasingly dependent on built-in self-test (BiST) options, embedded monitors and sensors, and redundancy and repair mechanisms to ensure higher interconnect yield in manufacturing and in field use.<\/p>\n<p>\u201cAny signal integrity issues after assembly and multi-die packaging become more difficult to diagnose and debug since probing is not easy,\u201d said Goriawalla. \u201cIn addition to the lane repair capabilities built into the HBM protocol itself, Synopsys offers SLM ext-RAM IP that provides built-in redundancy analysis and post-package repair to test the thousands upon thousands of interconnects, in HBM. During any service downtime mode, the user can run JEDEC-recommended algorithms via SLM ext-RAM to perform diagnosis and repair. This also enables a proactive approach. For instance, you may realize that certain failures occur in the field due to phenomena like aging or perhaps some marginalities.You don\u2019t want your LLM, which can take days or weeks to run, to fail because of this marginality. So you proactively swap out a marginal lane with a good lane, enabling that in-field, in-system.\u201d<\/p>\n<p>In-system testing and in-field diagnosis are extremely attractive to hyperscalar customers that are constantly pushing their systems for the highest uptime possible. \u201cDuring mission mode testing you are determining whether the IC is performing its function, as expected,\u201d said Goriawalla. \u201cAs it does the inference or the training, you use a contactless embedded monitoring system to see among all these lanes in the for die-to-die interfaces, if, for example, certain marginalities are occurring, or is the eye of the PHY getting smaller? Maybe you have a \u2018walking wounded\u2019 interconnect, and instead of letting it go to failure, at the next scheduled downtime, you take that marginal lane offline and swap it for a good lane.\u201d<\/p>\n<p>In-system tests use embedded deterministic test (EDT) patterns to enable targeted in-field testing to detect latent defects that may arise during any stage of the device lifecycle. During use, chip manufacturers increasingly need to account for the effects of thermal stress, workload-induced degradation, voltage fluctuations, and other causes of aging. For that reason, in-system test, once restricted to automotive and mission-critical systems, has found its way into data centers.<\/p>\n<p>Package-level interface reliability<br \/>Verifying the connectivity and functionality of XPU-XPU and XPU-HBM interfaces depends on well-established boundary scan testing, memory BiST, and at-speed functional testing. \u201cBoundary scan, particularly 1149.1 and 1149.6, serves as a workhorse for testing interconnects at both the board and package levels. Each chip, including the xPU and HBM, incorporates a boundary scan register around its I\/O pins. This method is valuable for detecting opens, shorts, and stuck-at faults on the interconnects without needing to fully operate the core logic,\u201d said Siemens EDA\u2019s Phan. \u201cFor high-speed, differential AC-coupled interfaces, common in xPU-xPU links, the 1149.6 standard is specifically designed to test their integrity, including detecting shorts between differential pairs.<\/p>\n<p>Phan noted that for HBM, dedicated MBiST or custom MBiST can be implemented within the SoC die. This MBIST generates specific data patterns, drives them across the HBM interface, and then reads them back from the HBM dies, thereby verifying the entire data path, including interposer traces, microbumps, and HBM I\/O logic. The HBM dies themselves often feature a loopback test mode that users can utilize to build a BiST for die-to-die interconnect tests at-speed. HBM also supports lane repair capability through the standard 1500 interfaces when a faulty lane is detected.<\/p>\n<p>\u201cFor high-speed serial links between xPUs (like PCIe, CXL, or proprietary interconnects), specialized SerDes (serializer\/deserializer) BiST is common. This involves activating loopback test modes (either internally or externally via package\/interposer traces), and pseudo-random binary sequence (PRBS) generation and checking,\u201d said Phan.<\/p>\n<p>Functional at-speed tests are essential when verifying connectivity. \u201cThese tests involve initiating large data transfers between the xPUs and HBM, or between two xPUs, and meticulously verifying the integrity of the data,\u201d Phan said. \u201cThis approach tests the entire communication stack, encompassing protocols and error correction mechanisms, ensuring that the interfaces perform as expected under real-world data loads.\u201d<\/p>\n<p>At the same time, Phan emphasized the increasing importance of I\/O or lane repair capabilities. \u201cThese features prevent the need to discard an entire chip or package due to localized defects,\u201d he said. \u201cThis built-in redundancy is critical for maintaining signal integrity and reducing manufacturing waste in complex AI accelerator systems.\u201d<\/p>\n<p>Long-term reliability and silicon lifecycle management are also important considerations in 2.5D\/3D chiplet-based packages. \u201cAging and stress-induced degradation require closer coordination between hardware and system software, with continuous tracking of parameters such as delay shifts, process monitoring, temperature, frequency, and eye width to enable timely decisions throughout the chip\u2019s lifecycle,\u201d said Surbhi Bansal, engineering director of DFT at Cadence. \u201cAs a result, DFT is evolving beyond traditional structural test. It now incorporates cross-die observability, such as in-situ eye-width measurement from the PHY during calibration and analog monitoring through ADC-based ATEST paths with digital readout, along with stress-aware test patterns. In addition, built-in redundancy and repair mechanisms across dies are part of the JEDEC spec and supported by our HBM PHY. Together, these capabilities enable more comprehensive visibility and resilience across the full lifecycle of the HBM stack.\u201d<\/p>\n<p>Failures in 2.5D\/3D architectures, TSV defectivity<br \/>As HBM stacks grow from 8 chiplets to 12, 16, and beyond, problems that once were mere nuisances are becoming major challenges. The rigors of assembly lead to warpage issues, which precipitate in cracks or misalignment of interconnects. \u201cIssues such as cracks (including latent defects), thermal distribution, and hotspots have already proved important, and these effects may become even more pronounced with higher stack density,\u201d said Advantest\u2019s Yokoyama. In terms of DFT, \u201capproaches such as programmable MBiST with more complex internal pattern execution may be considered to improve coverage.\u201d<\/p>\n<p>A major source of defectivity is in the through-silicon vias (TSVs). With each new generation of HBM, TSVs become more closely spaced, introducing more potential for failure. TSVs are lined with a thin dielectric barrier before the copper is deposited inside the vias, and any discontinuities in this layer can prove detrimental to yield. As the pitch of TSVs shrinks, so does the pitch of the underlying microbumps that connect each DRAM to the DRAM below it in the stack.<\/p>\n<p>HBM manufacturers typically use the via middle integration flow for TSV formation, meaning the TSVs are formed after the front-end processes (transistors), but before back-end metallization. The advantage to this approach is that the TSV can be integrated before the wafer is thinned, while it is still thick and mechanically robust. Forming TSVs at this juncture also avoids exposing the finished BEOL stack to the aggressive via etch, high-temperature oxide line deposition, and copper annealing steps. TSVs inside HBM are generally 2 to 5 microns in diameter and 30 to 60 microns deep (i.e., the depth of the wafer).<\/p>\n<p>When larger TSVs are exposed at the wafer surface, they may be contacted by a probe. However, given the tiny size of the TSV in HBM relative to probe needles (around 35\u00b5m), only TSV test structures can be contacted. Typically, hundreds or thousands of vias are connected in a daisy chain, particularly after backside wafer thinning and TSV reveal, where defects can be introduced. Any deviation from a normal in-series resistance can indicate potential defects such as opens, cracks, incomplete metal fill, misaligned contact, etc. But an outlier daisy chain result alone will not highlight which TSV has failed.<\/p>\n<p>\u201cWith the traditional daisy chain, they chain many bonds together, so really you have a continuity check, a go\/no go test. There is a loss of information about the individual bond because it\u2019s drowned out in the noise of the entire chain,\u201d said Modus Test\u2019s Lewis. \u201cWhat we need to do in the test vehicle design, instead of just chaining large chains, is make a Kelvin connection and get up around each one of these bonds and measure the bond \u2014 each individual bond with a Kelvin connection \u2014 very precisely into the micro-ohm range, by distributing [test structures] across each layer of the substrate, each die-to-die interconnect. It\u2019s all about planning, test insertion of the circuit, and then making the measurements and getting sufficient data for it to be valuable. We need thousands of measurements per package to take that information and train the inspection models, as well.\u201d<\/p>\n<p>Part of the change in TSV testing has to do with increasing signal integrity (SI) and power integrity (PI) issues as TSVs get closer together. \u201cThe industry\u2019s shift in TSV testing is being driven by several factors,\u201d said Cadence\u2019s Bansal. \u201cHigher speeds result in SI\/PI dominating functionality. There\u2019s also a need for at-speed, margin-aware validation, software methods to plot the eye, and system-level at-speed loopback tests.\u201d She noted the need for more comprehensive LFSR polynomials to send more exhaustive seeds for testing SI\/PI and other defects within and across lanes.<\/p>\n<p>When defects are found, HBM relies heavily on redundancy and repair to improve yield. Testing of the memory array typically identifies faulty rows or columns, spare resources are identified, and the repair program is executed. \u201cMemory repair is a built-in feature of the HBM die itself, because each die contains memory BiST engines specifically tailored to run sophisticated memory test algorithms to thoroughly test the memory cells,\u201d said Phan. \u201cOnce a faulty memory location is identified (by the HBM\u2019s BiST or the SoC\u2019s MBiST), the user can program the failing addresses and channels into SOFT_REPAIR or HARD_REPAIR Write Data Registers (WDRs) via the IEEE-1500 interfaces. After programming, a memory repair operation can be initiated, allowing the HBM to reconfigure itself to bypass the faulty elements, improving yield and reliability.\u201d<\/p>\n<p>There\u2019s an increasing need for power-aware simulation because of the great number of high-speed channels operating simultaneously. \u201cPower integrity is a growing challenge,\u201d said Bansal. \u201cThe increasing number of high-speed channels introduces greater susceptibility to IR drop and simultaneous switching noise, driving the need for more robust design and power-aware simulation techniques, such as staggered activation to reduce initial inrush current.\u201d<\/p>\n<p>Another technology that can help is power-aware automated test program generation. \u201cPower-aware ATPG is often a necessity for AI applications due to their high-power demands,\u201d said Siemens EDA\u2019s Phan. \u201cIt can minimize power consumption during the critical scan shift and scan capture operations through a combination of hardware insertion and pattern generation techniques.\u201d When combined with the Streaming Scan Network (SSN), power-aware ATPG can smooth the power profile through staggered shift clocks, reducing test time and power usage significantly.<\/p>\n<p>High temperature during testing also contributes to failures. \u201cAI workloads utilize full HBM bandwidth on a continuous basis. Hence, heat generation is becoming increasingly common with HBM,\u201d said Bansal. \u201cFor instance, high toggle rates cause localized heating, resulting in timing shifts, IR drop, and increased leakage. During test, heat generation is an issue with a high ATPG toggle rate. However, this can also help screen the parts during test.\u201d<\/p>\n<p>Together, these failure modes call for silicon lifecycle monitoring through manufacturing and into the field to enable software repair whenever possible and hardware replacements only as needed. In general, there is a trend to use data center systems longer, which may only be possible when adequate in-field test and repair mechanisms are on board.<\/p>\n<p>Conclusion<br \/>As HBMs grow in stack height and complexity, so does the testing approach and design-for-test strategy. The industry is in the process of rolling out HBM4 devices with stack heights of 16 DRAMs, some of which will have custom logic base dies to customize the configuration for specific workloads. DFT too will be customized to meet specialized needs and address the many failure modes of stacked die, including TSV film discontinuities and unlanded contacts, thermally-induced failures, timing shifts, SI\/PI failures, die cracks, etc.<\/p>\n<p>An important role of DFT is to verify the connectivity and functionality of xPU-xPU and xPU-HBM interfaces using boundary scan testing, memory BiST, and at-speed functional testing. Embedded monitors zoom in on critical metrics like timing margin, temperature changes, and voltage droop, enabling big data analysis and potential tracing of failures to their root cause. Ultra-precise Kelvin measurements in appropriate test structures can help ensure the quality of individual bonds, whether they are hybrid-bonded chip-to-chip or thermocompression-bonded microbumps.<\/p>\n<p>What\u2019s clear is that a plethora of tools and DFT techniques are needed to fully test advanced interconnects in stacked chips moving forward, and HBM is the proving ground for these methods.<\/p>\n<p>Related Articles<br \/><a href=\"https:\/\/semiengineering.com\/multi-die-testing-in-the-field-must-build-on-established-test-methodologies\/\" rel=\"nofollow noopener\" target=\"_blank\">Multi-Die Testing In The Field Must Build On Established Test Methodologies<\/a><br \/>Having a device that works at time zero is no longer a guarantee of reliability over its lifetime.<\/p>\n<p><a href=\"https:\/\/semiengineering.com\/ai-accelerators-usher-in-new-era-of-semiconductor-test\/\" rel=\"nofollow noopener\" target=\"_blank\">AI Accelerators Usher In New Era For IC Test<\/a><br \/>The number and variety of test interfaces, coupled with increased packaging complexity, are adding a slew of new challenges.<\/p>\n<p><a href=\"https:\/\/semiengineering.com\/ai-accelerator-testing-depends-on-dft-innovations\/\" rel=\"nofollow noopener\" target=\"_blank\">AI Accelerator Testing Depends On DFT Innovations<\/a><br \/>Multi-die assemblies greatly increase the number of things that can go wrong, and the difficulty of finding them.<\/p>\n<p><a href=\"https:\/\/semiengineering.com\/hbm-shifts-testing-left-to-preserve-ai-chip-yield\/\" rel=\"nofollow noopener\" target=\"_blank\">HBM Shifts Testing Left To Preserve AI Chip Yield<\/a><br \/>Testing sooner and more often can improve quality and reduce scrap, but it\u2019s also more costly.<\/p>\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"Key Takeaways: DFT is increasingly critical for detecting defects in memory cells, TSVs, microbumps, and high-speed interfaces, while&hellip;\n","protected":false},"author":2,"featured_media":589762,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[249142,249143,81770,249144,28212,88476,30054,167162,61,60,249145,167166,81766,13156,80,249146,99485],"class_list":["post-589761","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-3d-ics","tag-advantest","tag-cadence","tag-design-for-test","tag-dft","tag-dram","tag-hbm","tag-hybrid-bonding","tag-ie","tag-ireland","tag-modus-test","tag-proteantecs","tag-siemens-eda","tag-synopsys","tag-technology","tag-tsvs","tag-yield"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/589761","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=589761"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/589761\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/589762"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=589761"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=589761"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=589761"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}