{"id":273571,"date":"2025-11-05T18:51:10","date_gmt":"2025-11-05T18:51:10","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/273571\/"},"modified":"2025-11-05T18:51:10","modified_gmt":"2025-11-05T18:51:10","slug":"microsoft-built-a-fake-marketplace-to-test-ai-agents-they-failed-in-surprising-ways","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/273571\/","title":{"rendered":"Microsoft built a fake marketplace to test AI agents \u2014 they failed in surprising ways"},"content":{"rendered":"<p id=\"speakable-summary\" class=\"wp-block-paragraph\">On Wednesday, researchers at Microsoft released a new simulation environment designed to test AI agents, along with new research showing that current agentic models may be vulnerable to manipulation. Conducted in collaboration with Arizona State University, the research raises new questions about how well AI agents will perform when working unsupervised \u2014 and how quickly AI companies can make good on promises of an agentic future.<\/p>\n<p class=\"wp-block-paragraph\">The simulation environment, dubbed the <a href=\"https:\/\/www.microsoft.com\/en-us\/research\/publication\/magentic-marketplace-an-open-source-environment-for-studying-agentic-markets\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">\u201cMagentic Marketplace\u201d<\/a> by Microsoft, is built as a synthetic platform for experimenting on AI agent behavior. A typical experiment might involve a customer-agent trying to order dinner according to a user\u2019s instructions, while agents representing various restaurants compete to win the order.<\/p>\n<p class=\"wp-block-paragraph\">The team\u2019s initial experiments included 100 separate customer-side agents interacting with 300 business-side agents. Because the source code for the marketplace is open source, it should be straightforward for other groups to adopt the code to run new experiments or reproduce findings.<\/p>\n<p class=\"wp-block-paragraph\">Ece Kamar, managing director of Microsoft Research\u2019s AI Frontiers Lab, says this kind of research will be critical to understanding the capabilities of AI agents. \u201cThere is really a question about how the world is going to change by having these agents collaborating and talking to each other and negotiating,\u201d said Kamar. \u201cWe want to understand these things deeply.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The initial research looked at a mix of leading models, including GPT-4o, GPT-5, and Gemini-2.5-Flash, and found some surprising weaknesses. In particular, the researchers found several techniques businesses could use to manipulate customer agents into buying their products. The researchers noticed a particular falloff in efficiency as a customer agent was given more options to choose from, overwhelming the attention space of the agent.<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe want these agents to help us with processing a lot of options,\u201d Kamar says. \u201cAnd we are seeing that the current models are actually getting really overwhelmed by having too many options.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The agents also ran into trouble when they were asked to collaborate toward a common goal, apparently unsure of which agent should play what role in the collaboration. Performance improved when the models were given more explicit instructions on how to collaborate, but the researchers still saw the models\u2019 inherent capabilities as in need of improvement.<\/p>\n<p>Techcrunch event<\/p>\n<p>\n\t\t\t\t\t\t\t\t\tSan Francisco<br \/>\n\t\t\t\t\t\t\t\t\t\t\t\t\t|<br \/>\n\t\t\t\t\t\t\t\t\t\t\t\t\tOctober 13-15, 2026\n\t\t\t\t\t\t\t<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe can instruct the models \u2014 like we can tell them, step by step,\u201d Kamar said. \u201cBut if we are inherently testing their collaboration capabilities, I would expect these models to have these capabilities by default.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"On Wednesday, researchers at Microsoft released a new simulation environment designed to test AI agents, along with new&hellip;\n","protected":false},"author":2,"featured_media":273572,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[45],"tags":[182,3298,181,507,1281,74],"class_list":{"0":"post-273571","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-artificial-intelligence","8":"tag-ai","9":"tag-ai-agents","10":"tag-artificial-intelligence","11":"tag-artificialintelligence","12":"tag-microsoft","13":"tag-technology"},"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/273571","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=273571"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/273571\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/273572"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=273571"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=273571"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=273571"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}