Integrating complementary evidence supports contextualized substitution assessment

Within the evaluated five-material dataset, literature-derived per-use GHG indicators and review-derived user evidence did not produce fully aligned material profiles. Materials with lower per-use GHG indicators did not necessarily achieve the highest user-experience feature or consumer-approval scores, while highly rated materials did not always show the lowest per-use GHG indicators. This lack of complete alignment indicates that no single evaluated criterion was sufficient to characterize the substitution options considered in this study.

The proposed decision-support workflow provided a structured approach for examining this lack of alignment by integrating two complementary evidence domains: literature-derived per-use GHG indicators and review-derived user evidence. The workflow does not replace process-based life-cycle assessment, controlled usability testing, or comprehensive sustainability assessment. Instead, it provides a transparent structure for comparing material alternatives within the selected criteria, harmonization assumptions, and scenario-specific weights. The review-derived evidence should therefore be interpreted as a proxy for reported user experience and consumer approval rather than as representative evidence of population-wide preferences.

The integrated results illustrate how combining climate-related and user-oriented evidence can reveal trade-offs that would remain less visible if either evidence domain were considered independently. This interpretation is restricted to the evaluated materials, the selected per-use GHG assumptions, the available review-derived evidence, and the predefined MCDA scenarios. Accordingly, the resulting rankings should be understood as contextual decision-support outputs rather than general judgments about the overall sustainability of the materials.

Review-derived user evidence contextualizes practical suitability

The review-derived analyses showed that practical product attributes differed across the evaluated straw materials. The lexical and sentiment analyses identified recurring expressions related to cleaning, comfort, sensory experience, dimensions, portability, durability, and duration of use. These patterns provide descriptive context for how users discussed the products, but they should not be interpreted as direct measures of adoption or as evidence that attributes caused higher or lower consumer approval.

The results nevertheless indicate that per-use GHG indicators alone do not capture all dimensions relevant to practical substitution. Materials with favourable climate-related indicators may still be associated with usability or sensory concerns, while materials with higher consumer-approval scores may not have the lowest per-use GHG indicators. Review-derived user evidence therefore provided a complementary proxy for reported product experience within the available Amazon review dataset. It does not replace representative consumer surveys, controlled usability testing, or direct observation of long-term product use.

These descriptive patterns are broadly consistent with previous studies showing that practical performance and consumer response are relevant when evaluating straw alternatives. Liang et al.15 reported that ease of straw insertion and fracture affected purchase willingness in controlled consumer testing, while Gutierrez et al.16 documented substantial loss of paper-straw compressive strength after immersion. Paliwoda et al.17 further showed that environmental indicators can influence purchasing decisions for lifestyle products. The present review-derived findings complement, but do not replace, such controlled testing and survey evidence.

The post hoc shallow decision-tree analysis further supported interpretation of the integrated MCDA outputs by summarizing how the per-use GHG indicator and user-experience feature score separated higher, intermediate, and lower GHG-prioritized integrated MCDA scores within the five-material decision matrix. Consumer approval contributed to the integrated target score but was not entered as a separate explanatory variable. The tree was used as an explanatory representation of the calculated MCDA results rather than as a predictive model of consumer behaviour, market adoption, or material performance. Its thresholds are specific to the evaluated materials, input values, normalization procedure, and GHG-prioritized scenario and should not be treated as general product-design, environmental, or usability benchmarks.

Decision-support implications

The principal value of the proposed workflow lies in supporting transparent comparison among substitution options when multiple evidence sources must be considered simultaneously. Within the present case study, the evaluated alternatives showed trade-offs among literature-derived per-use GHG indicators, user-experience features, and rating-based consumer approval. Considering only one criterion could therefore favour an option that performs well on that dimension while overlooking limitations identified through the other evaluated evidence streams.

In environmental-management and procurement contexts, the workflow may serve as an exploratory screening aid for structuring comparisons before more detailed assessment or implementation. It can help decision-makers identify how rankings change when greater emphasis is placed on climate-related performance, user-experience features, or consumer approval. However, the outputs should not be used as stand-alone procurement recommendations or policy prescriptions without context-specific data, stakeholder-defined priorities, and additional environmental and usability evidence.

The workflow also improves transparency by displaying the contribution of each criterion to the integrated MCDA score and by examining the sensitivity of rankings across predefined weighting scenarios. This makes the basis of the comparison more explicit than a single aggregated recommendation. The resulting rankings remain conditional on the selected materials, harmonized per-use GHG indicators, review-derived user evidence, normalization choices, and scenario-specific weights.

Although demonstrated using drinking-straw alternatives, the workflow may be adapted to other product-substitution contexts when comparable environmental and user-oriented evidence can be assembled. Such applications require product-specific criteria, appropriate evidence harmonization, clearly justified weights, and independent validation before operational use.

Study limitations and future directions

Several limitations should be considered when interpreting the findings. First, the environmental evidence was restricted to literature-derived per-use GHG indicators harmonized from previously published studies rather than generated through a new process-based life-cycle assessment. Although harmonization enabled comparison on a common per-use basis, differences in original functional units, system boundaries, reuse assumptions, washing conditions, transport, durability, and end-of-life scenarios could not be fully eliminated. The assessment therefore represents a climate-related indicator comparison and should not be interpreted as a comprehensive sustainability assessment.

Second, the analysis did not include other potentially relevant environmental dimensions, such as resource depletion, toxicity, water use, littering potential, persistence, recyclability, compostability, or leakage into terrestrial and marine environments. The relative rankings among materials may change if these dimensions are incorporated. Future studies should therefore extend the environmental evidence base using product-specific criteria and, where feasible, consistently modelled life-cycle inventories.

Third, the user-oriented evidence was derived from Amazon customer reviews and should be interpreted as review-derived proxy evidence rather than as a representative consumer survey or controlled usability assessment. Online reviews may be affected by self-selection, platform-specific user populations, product availability, reviewer expectations, and differences among individual products grouped within the same material category. The rating-based consumer-approval score and lexicon-based sentiment classifications also capture different aspects of the review evidence and should not be treated as interchangeable measures.

Fourth, the user-experience feature score was based on simplified binary coding of five usability and safety attributes. This representation facilitated transparent integration within the MCDA but did not capture the intensity, frequency, severity, or context of reported product attributes. Similarly, lexical frequencies describe recurring expressions within review-level sentiment subsets but do not establish the independent polarity, prevalence among all users, or causal effect of individual product characteristics.

Fifth, the integrated MCDA was based on a five-material decision matrix, min–max normalization, and four predefined weighting scenarios. The resulting scores and rankings are therefore conditional on the evaluated materials, selected per-use GHG assumptions, available review-derived evidence, normalization procedure, and scenario-specific weights. Weights were used to examine sensitivity to different decision priorities and should not be interpreted as universal sustainability preferences. Broader applications should incorporate stakeholder-defined weights, alternative normalization procedures, uncertainty and sensitivity analyses, and validation using additional product categories and datasets.

The shallow decision tree and Apriori association-rule analysis were constrained by the five-material decision matrix and simplified binary feature coding. Accordingly, the decision-tree thresholds and association rules should be interpreted only as internal explanatory and exploratory descriptive patterns, respectively, rather than as independently validated or externally generalizable predictive, causal, environmental, usability, product-design, or consumer-behaviour relationships.

Finally, conventional plastic straws were retained as a contextual baseline but were not included as a scored alternative in the integrated MCDA because comparable review-derived user-experience and consumer-approval evidence was not consistently available under the same inclusion criteria. Their exclusion limits direct comparison between conventional plastic and the evaluated alternatives. Future research could address this limitation by assembling fully harmonized environmental and user-oriented evidence for the baseline product and by expanding the analysis to include additional materials, platforms, user groups, and product-use contexts.

Future studies should also examine uncertainty propagation during evidence harmonization, test alternative MCDA methods, and evaluate the workflow prospectively in procurement, product-design, or policy contexts. Such extensions would help determine whether the workflow remains informative when applied to larger decision matrices and independently collected environmental and user-experience data.