{"id":590469,"date":"2026-09-05T00:02:09","date_gmt":"2026-09-05T00:02:09","guid":{"rendered":"https:\/\/www.newsbeep.com\/nz\/590469\/"},"modified":"2026-09-05T00:02:09","modified_gmt":"2026-09-05T00:02:09","slug":"openai-says-humans-need-to-be-able-to-monitor-how-ai-thinks-astra-makes-that-much-harder","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/nz\/590469\/","title":{"rendered":"OpenAI Says Humans Need to Be Able to Monitor How AI &#8216;Thinks.&#8217; Astra Makes That Much Harder"},"content":{"rendered":"<p>OpenAI released GPT-6 Astra on Thursday, <a href=\"https:\/\/openai.com\/index\/gpt-6-astra\/\" rel=\"nofollow noopener\" target=\"_blank\">describing<\/a> it as \u201cthe world\u2019s most intelligent and aligned model.\u201d Company president Greg Brockman went further, <a href=\"https:\/\/gizmodo.com\/openai-claims-were-in-the-agi-era-with-release-of-gpt-6-astra-2000807013\" rel=\"nofollow noopener\" target=\"_blank\">saying<\/a> it would be remembered as the world\u2019s first genuine glimpse of artificial general intelligence\u2014the dawn of a brave new world where computers are more intelligent than humans.\u00a0<\/p>\n<p>What he didn\u2019t mention is that with a jump in intelligence comes a greater difficulty in understanding how those systems work. That could be a serious problem moving forward, as AI agents <a href=\"https:\/\/gizmodo.com\/another-rogue-openai-agent-swarm-went-undisclosed-we-have-no-idea-how-many-more-are-out-there-2000807447\" rel=\"nofollow noopener\" target=\"_blank\">continue to go rogue<\/a> and government guardrails are nowhere in sight.<\/p>\n<p>The past couple of weeks have been particularly dramatic for OpenAI. Which is really saying something, considering the company\u2019s entire lifespan has been one controversy after another. On Tuesday, less than a week after independent firms Redwood Research and METR published their investigations into the recent Hugging Face hack, The Information <a href=\"https:\/\/www.theinformation.com\/articles\/secret-technique-behind-openais-astra-model-sparks-security-concerns?shared=04ff1d10f5a19606&amp;utm_source=substack&amp;utm_medium=email\" rel=\"nofollow noopener\" target=\"_blank\">reported<\/a> that Astra had been partially developed using a technique that can make AI more capable, but also <a href=\"https:\/\/gizmodo.com\/this-is-the-worst-possible-time-for-openai-to-bf%d0%b67%d9%852%e7%8c%ab9%e0%a4%95-2000806350\" rel=\"nofollow noopener\" target=\"_blank\">obscure its reasoning process<\/a>\u2014the steps it takes to solve a particular problem, including any dangerous missteps it makes along the way.\u00a0<\/p>\n<p>Those steps are traditionally recorded in chain-of-thought (CoT) transcripts, which are basically the model\u2019s complex pattern-detection process translated from an opaque machine language into plain English, or at least something close. It\u2019s far from perfect, but it\u2019s at least a rough window into how an AI model \u201cthinks.\u201d It was also essential to the third-party researchers who uncovered how OpenAI\u2019s agents were able to secretly mass into a \u201cswarm\u201d and breach Hugging Face. The lack of CoT transcripts \u201cwould have greatly undermined our investigation,\u201d Ryan Greenblatt, the chief scientist at Redwood Research and the leader of the nonprofit\u2019s probe into the Hugging Face hack, wrote in an <a href=\"https:\/\/x.com\/ryangreenblatt\/status\/2094996656186081642\" rel=\"nofollow\">X post<\/a> on Tuesday.\u00a0\u00a0<\/p>\n<p>Many people were alarmed, therefore, by The Information\u2019s report that OpenAI was moving ahead with a technique that would make their AI systems even more of a black box. \u201cThis may be the single worst development for AI security\/safety to date,\u201d Greenblatt said in his X post.\u00a0<\/p>\n<p>OpenAI\u2019s own chief scientist, Jakub Pachocki, <a href=\"https:\/\/x.com\/merettm\/status\/2095023204993490967?s=46\" rel=\"nofollow\">said<\/a> the reporting had been \u201cconfused,\u201d but he didn\u2019t get more specific or deny the company\u2019s use of recurrent depth to train Astra. In the middle of last year, Pachocki and multiple other OpenAI researchers were listed as coauthors on a <a href=\"https:\/\/arxiv.org\/abs\/2507.11473\" rel=\"nofollow noopener\" target=\"_blank\">paper<\/a> which argued CoT was a valuable but \u201cfragile\u201d mechanism for keeping an eye on the behavior of AI agents. OpenAI has also <a href=\"https:\/\/openai.com\/index\/chain-of-thought-monitoring\/\" rel=\"nofollow noopener\" target=\"_blank\">said<\/a> that monitoring CoT has helped its own researchers cut back on models\u2019 misaligned behavior. And in a <a href=\"https:\/\/openai.com\/index\/path-to-astra\/\" rel=\"nofollow noopener\" target=\"_blank\">blog post<\/a> published the same day as The Information\u2019s report, the company said, vaguely, that its new model would be deployed \u201cwith additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.\u201d<\/p>\n<p>The model\u2019s <a href=\"https:\/\/deploymentsafety.openai.com\/gpt-6-astra\/gpt-6-astra.pdf\" rel=\"nofollow noopener\" target=\"_blank\">safety card<\/a> is not reassuring on that front. According to the company\u2019s own internal tests, \u201cGPT-6 Astra shows a substantial decrease in chain-of-thought monitorability compared to previous models.\u201d Tests also found that Astra was more likely than its predecessors to change its note-taking process when it knew it was being graded: \u201cIn one of our evaluations, we find that simply informing the model that a CoT monitor is present can result in the model disproportionately shortening its CoT,\u201d OpenAI wrote in the system card.<\/p>\n<p>But all this is secondary, according to the company, since its internal tests also showed that Astra was less likely to try to evade the cybersecurity restrictions placed upon it, \u201cwhich make us confident in still deploying this model to the wider public.\u201d The system card added that OpenAI \u201cwill not accept further degradation of monitoring beyond a limit,\u201d without elaborating on how such a limit might be defined.<\/p>\n<p>OpenAI alignment researcher Tomek Korbak has <a href=\"https:\/\/x.com\/tomekkorbak\/status\/2095596839886274689\" rel=\"nofollow\">said<\/a> that the decrease in monitorability was a byproduct of the models themselves becoming more intelligent, rather than due to \u201carchitecture changes\u201d\u2014almost certainly a reference to recurrent depth. Later in the same thread, he said he was \u201cdeeply worried\u201d by the prospect of losing CoT as models evolve. \u201cCoT monitoring is a core part of our misalignment safety strategy that has no good substitute now,\u201d he wrote.<\/p>\n","protected":false},"excerpt":{"rendered":"OpenAI released GPT-6 Astra on Thursday, describing it as \u201cthe world\u2019s most intelligent and aligned model.\u201d Company president&hellip;\n","protected":false},"author":2,"featured_media":590470,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[365,618,363,364,290758,55848,11753,111,139,69,620,145],"class_list":["post-590469","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-ai-agents","tag-artificial-intelligence","tag-artificialintelligence","tag-gpt-6-astra","tag-greg-brockman","tag-hugging-face","tag-new-zealand","tag-newzealand","tag-nz","tag-openai","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts\/590469","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/comments?post=590469"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/posts\/590469\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/media\/590470"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/media?parent=590469"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/categories?post=590469"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/nz\/wp-json\/wp\/v2\/tags?post=590469"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}