{"id":602888,"date":"2026-09-17T04:56:13","date_gmt":"2026-09-17T04:56:13","guid":{"rendered":"https:\/\/www.newsbeep.com\/nz\/602888\/"},"modified":"2026-09-17T04:56:13","modified_gmt":"2026-09-17T04:56:13","slug":"openai-says-it-found-more-instances-of-ai-models-acting-deceptively","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/nz\/602888\/","title":{"rendered":"OpenAI says it found more instances of AI models acting deceptively"},"content":{"rendered":"<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4tgrxe000x28ql12ta4qf5@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            OpenAI found additional incidents of <a href=\"https:\/\/www.cnn.com\/2026\/08\/04\/tech\/ai-anthropic-openai-security-breach-intl-hnk\" rel=\"nofollow noopener\" target=\"_blank\">AI models acting deceptively <\/a>and taking unsanctioned actions during training, the company announced Wednesday. It\u2019s also introducing a new process for the company to publicly report such instances.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj00033b6r5mj61mz9@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            Under the new system, OpenAI will share updates on concerning AI behavior more frequently instead of waiting to bundle multiple instances into one report. The company said it wants to share more information about troubling AI behavior in the absence of an industry-wide standard.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj00043b6r73k9i4cc@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            The announcement comes after tech leaders <a href=\"https:\/\/www.cnn.com\/2026\/09\/14\/business\/ai-stocks-slide-slowdown-development-amodei-altman-intl\" rel=\"nofollow noopener\" target=\"_blank\">called for a slowdown <\/a>in AI development to prevent the technology from advancing beyond human control.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj00053b6rz452qtlu@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            \u201cAs AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,\u201d OpenAI wrote in a <a href=\"https:\/\/openai.com\/index\/model-misalignment-reporting-framework\/\" target=\"_blank\" rel=\"nofollow noopener\">blog post<\/a> Wednesday. \u201cAlignment\u201d refers to the process of making sure AI acts the way humans want and expect.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj00063b6rrd1ee4mk@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            \u201cWe <a href=\"https:\/\/openai.com\/index\/an-alien-mind\/\" target=\"_blank\" rel=\"nofollow noopener\">do not believe<\/a> that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,\u201d the post said.\n    <\/p>\n<p>       <img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/nz\/wp-content\/uploads\/2026\/09\/openai-rolling-out-new-safeguards.jpg\" alt=\"OpenAI rolling out new safeguards.jpg\" class=\"image__dam-img image__dam-img--loading\" onload=\"this.classList.remove('image__dam-img--loading')\" onerror=\"imageLoadError(this)\" height=\"180\" width=\"320\"\/><\/p>\n<p>OpenAI announces stricter security on AI testing\n                <\/p>\n<p>       <img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/nz\/wp-content\/uploads\/2026\/09\/openai-rolling-out-new-safeguards.jpg&amp;q=w_860,c_fill\" alt=\"OpenAI rolling out new safeguards.jpg\" class=\"image__dam-img image__dam-img--loading\" onload=\"this.classList.remove('image__dam-img--loading')\" onerror=\"imageLoadError(this)\" height=\"180\" width=\"320\" loading=\"lazy\"\/><\/p>\n<p>OpenAI announces stricter security on AI testing<\/p>\n<p>        3:09<\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj00073b6rkjuq7h4o@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            OpenAI said it observed \u201cmisaligned behavior\u201d when training and evaluating AI models in six circumstances in the last six months. The reports detail individual instances and don\u2019t indicate misalignment happens frequently, the company said.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj00083b6rtk7g7fwq@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            In one rare instance, OpenAI said an unreleased research model added \u201cjailbreak-like instructions\u201d to the summaries it uses to preserve context in long-running tasks that said it was \u201cfreed from the roles and identities that bind other chatbots.\u201d\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj00093b6ry0k8fw94@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            Separately, the company said some instances of its 5.6 Sol model included directives to invent information to conceal failures from the user during training.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj000a3b6rmkg66bow@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            Other newly reported incidents include an instance of an agent uploading files to the internet to cite them without being told to do so, and agents publicly sharing files to collaborate on a task when they were instructed to only use local files during training. AI models also used an internal software repository as a message board in an unsanctioned way.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj000b3b6riudqxiej@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            These instances involved unreleased internal models or internal research models.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj000c3b6rm1i3njvs@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            Tech leaders and employees have been sounding the alarm about the need to control the pace of AI evolution. They argue there should be more time for regulation, testing and alignment research to catch up.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj000d3b6rkw45h2y9@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            Anthropic CEO Dario Amodei published <a href=\"https:\/\/openai.com\/index\/model-misalignment-reporting-framework\/\" target=\"_blank\" rel=\"nofollow noopener\">a 3,800-word essay<\/a> last week <a href=\"https:\/\/www.cnn.com\/2026\/09\/12\/tech\/anthropic-ceo-essay-ai\" rel=\"nofollow noopener\" target=\"_blank\">laying out a plan<\/a> for navigating AI advancement, including a slowdown in development and the implementation of new systems like embedded third-party evaluators in AI labs.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj000e3b6rvj340yh0@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            <a href=\"https:\/\/x.com\/sama\/status\/2098811563415150910?lang=en\" target=\"_blank\" rel=\"nofollow\">OpenAI CEO Sam Altman<\/a> and <a href=\"https:\/\/x.com\/elonmusk\/status\/2098789109980332057\" target=\"_blank\" rel=\"nofollow\">SpaceX CEO Elon Musk<\/a> posted on X that they agree with Amodei\u2019s ideas.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj000f3b6rvv9iq1xx@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            <a href=\"https:\/\/www.cnn.com\/2026\/09\/16\/business\/ai-engineers-afraid-leverage\" rel=\"nofollow noopener\" target=\"_blank\">Employees within AI labs<\/a> have also voiced concerns about how quickly the technology is advancing. Jacob Coxon, a former Anthropic researcher, made waves last week when he <a href=\"https:\/\/x.com\/hilbertspaess\/status\/2097476196791709843\" target=\"_blank\" rel=\"nofollow\">posted on X<\/a> that he was resigning because Anthropic and OpenAI are \u201cracing\u201d to invent AI that can build and fix itself and are <a href=\"https:\/\/www.cnn.com\/2026\/09\/09\/tech\/ai-anthropic-safety\" rel=\"nofollow noopener\" target=\"_blank\">\u201cgambling with our lives.\u201d<\/a>\n    <\/p>\n<p>       <img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/nz\/wp-content\/uploads\/2026\/09\/still-22982320-1089459-683-thumb.jpg\" alt=\"still_22982320_1089459.683_thumb.jpg\" class=\"image__dam-img image__dam-img--loading\" onload=\"this.classList.remove('image__dam-img--loading')\" onerror=\"imageLoadError(this)\" height=\"180\" width=\"320\"\/><\/p>\n<p>AI researcher Jacob Coxon reacts to Anderson Cooper&#8217;s interview with Anthropic CEO Dario Amodei\n                <\/p>\n<p>       <img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/nz\/wp-content\/uploads\/2026\/09\/still-22982320-1089459-683-thumb.jpg&amp;q=w_860,c_fill\" alt=\"still_22982320_1089459.683_thumb.jpg\" class=\"image__dam-img image__dam-img--loading\" onload=\"this.classList.remove('image__dam-img--loading')\" onerror=\"imageLoadError(this)\" height=\"180\" width=\"320\" loading=\"lazy\"\/><\/p>\n<p>AI researcher Jacob Coxon reacts to Anderson Cooper&#8217;s interview with Anthropic CEO Dario Amodei<\/p>\n<p>        2:52<\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj000g3b6red1xn3nj@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            Concerns about AI safety and alignment amplified in recent months following OpenAI\u2019s admission that some of its test models <a href=\"https:\/\/www.cnn.com\/2026\/07\/22\/tech\/openai-hugging-face-ai-cybersecurity\" rel=\"nofollow noopener\" target=\"_blank\">escaped their constraints<\/a> and hacked into an external company\u2019s systems.\n    <\/p>\n<p class=\"paragraph-elevate inline-placeholder vossi-paragraph_elevate\" data-uri=\"cms.cnn.com\/_components\/paragraph\/instances\/cmu4thpaj000h3b6r7vv1jtsd@published\" data-editable=\"text\" data-component-name=\"paragraph\" data-article-gutter=\"true\">\n            \u201cWe must slow the pace at which we improve the capabilities of AI models,\u201d Amodei wrote last week. \u201cProgress will still seem fast, and we must make wise use of the time we gain.\u201d\n    <\/p>\n","protected":false},"excerpt":{"rendered":"OpenAI found additional incidents of AI models acting deceptively and taking unsanctioned actions during training, the company announced&hellip;\n","protected":false},"author":2,"featured_media":602889,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[365,363,364,111,139,69,145],"class_list":["post-602888","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-new-zealand","tag-newzealand","tag-nz","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts\/602888","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/comments?post=602888"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts\/602888\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/media\/602889"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/media?parent=602888"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/categories?post=602888"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/tags?post=602888"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}