{"id":439256,"date":"2026-05-13T07:45:16","date_gmt":"2026-05-13T07:45:16","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/439256\/"},"modified":"2026-05-13T07:45:16","modified_gmt":"2026-05-13T07:45:16","slug":"accelerating-detection-engineering-using-ai-assisted-synthetic-attack-logs-generation","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/439256\/","title":{"rendered":"Accelerating detection engineering using AI-assisted synthetic attack logs generation"},"content":{"rendered":"<p>\t\tIn this article<\/p>\n<p class=\"wp-block-paragraph\">Logs and telemetry\u00a0are\u00a0the\u00a0foundation\u00a0of modern cybersecurity. They enable threat detection, incident response, forensic investigation, and\u00a0compliance across endpoints, networks, and cloud environments. Yet, despite their importance, high\u2011quality security\u00a0attack\u00a0logs are notoriously difficult to\u00a0collect,\u00a0especially\u00a0at scale.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Real\u2011world security\u00a0telemetry is\u00a0often composed of repeated benign activity occurring\u00a0across\u00a0environments\u00a0and\u00a0with\u00a0very\u00a0rare\u00a0malicious activity.\u00a0Gathering, labeling, and\u00a0maintaining\u00a0datasets with real attack logs is costly and operationally challenging. It requires not only labeling malicious activities, but also fully reconstructing attack scenarios.\u00a0These challenges significantly slow\u00a0detection engineering and limit\u00a0the\u00a0quality of both the rule-based detection\u00a0authoring\u00a0and\u00a0anomaly-detection\u00a0approaches.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In this post, we explore a different path:\u00a0using\u00a0AI\u00a0to generate realistic, high\u2011fidelity synthetic security\u00a0attack\u00a0logs. By translating attacker behaviors,\u00a0expressed as tactics, techniques, and procedures\u00a0(TTPs)\u2014directly into structured telemetry, we aim to accelerate detection development while preserving realism and security.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Why is this work important for\u00a0Microsoft\u00a0Defender customers?\u00a0<\/p>\n<p class=\"wp-block-paragraph\">For Microsoft Defender customers, this work is crucial because\u00a0it directly addresses the challenge of obtaining high-quality, realistic security attack logs needed for effective threat detection and response. By\u00a0leveraging\u00a0AI-driven synthetic log generation, organizations can accelerate the development of detection rules\u00a0and AI-based automation approaches,\u00a0while\u00a0ensuring\u00a0privacy and reducing operational overhead. Synthetic logs enable customers to simulate a broader range of attack scenarios\u2014including rare and emerging threats\u2014without exposing sensitive data or relying on costly lab-based simulations.\u00a0Ultimately, this\u00a0approach enhances the agility and effectiveness of\u00a0Microsoft Defender detection and response capabilities,\u00a0helping customers stay ahead of evolving cyber threats.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Why Synthetic Security Logs\u00a0in addition to\u00a0Lab Simulations?\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Synthetic data has been widely adopted in various fields as a privacy-conscious substitute for real data, and it offers even greater advantages in cybersecurity. It enables the creation of safe, shareable datasets that avoid exposure of sensitive customer information, allows simulation of rare or emerging attacks that are challenging to\u00a0observe\u00a0in real environments, accelerates the process of detection engineering and testing, and supports reproducible experiments for benchmarking and evaluation.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">While synthetic logs are not a replacement for all lab-based validation, they can complement lab simulations by speeding up early-stage detection design, testing, and coverage expansion.\u00a0Traditionally, generating realistic attack telemetry\u00a0requires\u00a0executing real attacks in controlled lab environments. While\u00a0accurate, this approach is slow, labor\u2011intensive, and difficult to scale.\u00a0It also limits agility for the security teams responsible for defending our systems and delays the rollout of new threat detections into production.\u00a0This\u00a0blog examines\u00a0whether\u00a0AI-assisted\u00a0synthetic log generation\u00a0can provide similar fidelity,\u00a0without the operational overhead of lab\u2011based attack execution.\u00a0<\/p>\n<p>Core Idea: From TTPs to Logs<\/p>\n<p class=\"wp-block-paragraph\">Attackers can abuse TTP through various actions that exploit different processes.\u00a0At\u00a0a high level,\u00a0the\u00a0proposed\u00a0workflow\u00a0consumes\u00a0\u201cTTP + Action\u201d\u00a0as input and produces\u00a0structured security logs\u00a0as output.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Input:\u00a0High\u2011level\u00a0attacker TTPs from\u00a0the\u00a0MITRE ATT&amp;CK\u00a0framework\u00a0[1], a widely used knowledge base of adversary tactics and techniques, and concrete attacker actions.\u00a0See the example\u00a0below.\u00a0<\/p>\n<p>Tactic\u00a0Technique\u00a0Action\u00a0Stealth\u00a0T1202 \u2013 Indirect Command Execution\u00a0\u00a0The attackers executed\u00a0forfiles\u00a0and obfuscated their actions using variable expansion of\u00a0%PROGRAMFILES\u00a0and hex characters (for example, 0x5d). They obfuscated the use of\u00a0echo, open, read, find,\u00a0and exec to extract file contents, then passed the output to a Python interpreter for execution.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Output:\u00a0Realistic log entries with correctly populated fields such as\u00a0\u201cCommand Line\u201d, \u201cProcess Name\u201d, \u201cParent Process Name\u201d,\u00a0and other relevant telemetry fields.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Goal: The goal is not to reproduce logs verbatim, but to generate\u00a0realistic, semantically correct logs\u00a0that\u00a0would accurately\u00a0trigger\u00a0detections,\u00a0mirroring real attacker behavior.\u00a0<\/p>\n<p>Approaches\u00a0for\u00a0Synthetic\u00a0Attack\u00a0Log Generation<\/p>\n<p class=\"wp-block-paragraph\">We\u00a0explore\u00a0three increasingly sophisticated techniques for generating logs.\u00a0<\/p>\n<p>Prompt\u2011Engineered Generation:\u00a0Our baseline approach uses a\u00a0series of\u00a0carefully designed\u00a0expert\u2011crafted prompts.\u00a0The\u00a0workflow\u00a0comprises\u00a0a structured, multi\u2011stage dialogue:\u00a0<\/p>\n<p>Prompting: The model is given a detailed attack scenario and context.\u00a0<\/p>\n<p>Iterative Generation: Logs are generated across multiple turns to\u00a0maintain\u00a0coherence.\u00a0<\/p>\n<p>Evaluation: An independent\u00a0large language model (LLM)-as-a-Judge\u00a0assesses realism and consistency.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">As depicted in\u00a0the following image, the prompts explicitly instruct the model to reason like a cybersecurity researcher, leverage MITRE\u00a0ATT&amp;CK knowledge, and produce coherent attack narratives.\u00a0<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-59.webp\" alt=\"\" class=\"wp-image-147322 webp-format\"  data-orig-src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-59.webp\"\/>Diagram that shows a three-stage AI agent pipeline: prompting for attack scenarios,<br \/>iterative generation of logs, and LLM-as-a-Judge evaluation.<\/p>\n<p>Agentic\u00a0Workflow-based Generation:\u00a0While the first approach works well in simpler cases, it struggles with complex, multi\u2011stage scenarios.\u00a0To address these limitations, we introduced an\u00a0agentic workflow\u00a0using three specialized agents\u00a0focused on different tasks:\u00a0<\/p>\n<p>Generator Agent: Produces\u00a0an initial\u00a0set of logs based on the input.\u00a0<\/p>\n<p>Evaluator Agent: Reviews logs and provides structured feedback.\u00a0<\/p>\n<p>Improver Agent: Suggests targeted refinements based on feedback.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">As depicted in the image below, these agents collaborate in an iterative loop\u00a0(generate, evaluate, improve),\u00a0allowing the system to correct errors, fill gaps, and refine details over multiple turns.\u00a0This collaborative process significantly improves log completeness and fidelity, especially for complex attack chains.\u00a0<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-60.webp\" alt=\"\" class=\"wp-image-147323 webp-format\"  data-orig-src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-60.webp\"\/>Diagram that shows a cyclical agentic workflow where generator, evaluator, and improver<br \/>agents collaborate to produce synthetic telemetry logs.<\/p>\n<p>Multi-Turn Reinforcement Learning with Verifiable Rewards:\u00a0While the synthetic logs generated by the agentic workflow are often semantically correct,\u00a0preserving key properties like parent\u2011child process relationships and event ordering,\u00a0they still differ noticeably from real event logs, especially in process paths,\u00a0command\u2011line arguments, service names and so on.\u00a0This limits the usage of these logs to test detection efficacy;\u00a0effective detection engineering\u00a0requires\u00a0reliably distinguishing\u00a0benign\u00a0activity\u00a0from malicious behavior.\u00a0\u00a0<br \/>To address this challenge,\u00a0we\u00a0conduct experiments using\u00a0Reinforcement Learning with Verifiable Rewards (RLVR).\u00a0Instead of rigid rewards used by the evaluator agent in the\u00a0previous\u00a0agentic workflow approach,\u00a0we use\u00a0partial\u00a0rewards\u00a0to learn the policies as follows:\u00a0<\/p>\n<p>We use an LLM\u2011as\u2011a\u2011Judge as follows to compare the synthesized data against ground\u2011truth logs.\u00a0\u00a0<\/p>\n<p>The model\u00a0only\u00a0awards\u00a0partial rewards based on semantic alignment\u00a0and\u00a0imposes a\u00a0penalty\u00a0if the\u00a0generated string is\u00a0not\u00a0an\u00a0exact\u00a0match of the ground-truth logs, producing a more context-aware and flexible reward signal to guide the learning process.\u00a0<\/p>\n<p>The judge also produces reasoning, making\u00a0evaluations\u00a0transparent,\u00a0and auditable.\u00a0<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-61.webp\" alt=\"\" class=\"wp-image-147324 webp-format\"  data-orig-src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-61.webp\"\/>Diagram that shows the LLM-as-a-Judge evaluation comparing generated logs to ground<br \/>truth, issuing rewards or penalties to drive policy updates.<\/p>\n<p class=\"wp-block-paragraph\">While this direction of research shows a lot of promise, it is heavily dependent on the amount of labeled training data.\u00a0To\u00a0address this limitation, we applied\u00a0data augmentations, including:\u00a0<\/p>\n<p>Paraphrasing attack narratives while preserving technical intent\u00a0<\/p>\n<p>Perturbing parameters (e.g., replacing executable names with plausible alternatives, re-ordering flags, etc.)\u00a0<\/p>\n<p class=\"wp-block-paragraph\">This allowed us to scale from hundreds to thousands of training examples.\u00a0<\/p>\n<p>Evaluation\u00a0Datasets<\/p>\n<p class=\"wp-block-paragraph\">To ensure our approach generalizes across environments and attack types, we evaluated it on three complementary datasets:\u00a0<\/p>\n<p>Goal\u2011Driven (GD) Campaigns: These are tightly scoped\u00a0datasets produced by\u00a0repeatable attack simulations\u00a0conducted by our threat researchers.\u00a0GDs are\u00a0built around a specific\u00a0security\u00a0objective\u00a0(e.g., detecting credential dumping on Windows servers). They provide clean ground truth and well\u2011defined attacker actions.\u00a0We used a total of 10 different GD executions to evaluate our approaches.\u00a0<\/p>\n<p>Security Datasets Project:\u00a0An open\u2011source initiative\u00a0[2]\u00a0that provides malicious and benign datasets from multiple platforms, enabling broader evaluation and generalizability across different environments.\u00a0\u00a0<\/p>\n<p>ATLASv2 Dataset:\u00a0The ATLASv2\u00a0dataset\u00a0[3]\u00a0is\u00a0comprised\u00a0of\u00a0Windows Security Auditing logs, Sysmon logs, Firefox logs, and\u00a0Domain Name System (DNS)\u00a0telemetry. These\u00a0logs\u00a0are\u00a0generated\u00a0across\u00a0two Windows VMs\u00a0by executing\u00a010\u00a0multi\u2011stage attack scenarios and introducing\u00a0realistic noise and cross\u2011host behaviors.\u00a0We limited the evaluation of synthetic attack logs to\u00a0malicious\u00a0activity\u00a0during the attack windows.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Note: The external datasets from the Security Datasets Project and ATLASv2 are used strictly for research and validation of our log generation methods. These datasets are not used in the development, training, or deployment of any commercial products.\u00a0<\/p>\n<p>Evaluation\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Methodology:\u00a0We evaluated\u00a0the prompt engineering and agentic workflow approach\u00a0on the three datasets\u00a0across multiple reasoning and non\u2011reasoning models, using\u00a0recall\u00a0as our primary metric.\u00a0Recall measures the model\u2019s ability to generate\u00a0semantically\u00a0relevant log instances (true positives) expected for a given attack scenario.\u00a0Our\u00a0LLM\u2011as\u2011a\u2011Judge\u00a0performs flexible matching, focusing on:\u00a0<\/p>\n<p class=\"wp-block-paragraph\">For example, a synthetic log\u00a0containing\u00a0\u201cforfiles.exe\u201d\u00a0can successfully match a ground\u2011truth entry with the full path\u00a0\u201cD:\\Windows\\System32\\forfiles.exe\u201d.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Key Results:\u00a0The results\u00a0in experimental evaluation\u00a0demonstrate\u00a0that prompt-only\u00a0\u00a0approaches\u00a0establish a baseline but show inconsistent performance.\u00a0The agentic workflows deliver dramatic recall improvements across all datasets.\u00a0Reasoning models, combined with agentic refinement, achieve the highest fidelity.\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Finally, our\u00a0experiments\u00a0training reinforcement learning\u00a0approaches\u00a0conclude\u00a0that while it shows a significant promise, a\u00a0substantial\u00a0amount of labeled data will be\u00a0required\u00a0for the agent to learn effective policies to make the synthetic data identical to benign logs.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Table 1 and Table 2\u00a0report\u00a0the performance of the prompt-based and agentic workflow-based approaches, respectively.\u00a0For reasoning\u00a0models\u00a0(o1, o3 and o3-mini), we report the recall values using\u00a0a\u00a0Medium\u00a0reasoning effort.\u00a0Overall, agentic collaboration\u00a0emerges\u00a0as the most effective technique for high\u2011quality synthetic attack logs generation.\u00a0<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-62.webp\" alt=\"\" class=\"wp-image-147325 webp-format\"  data-orig-src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-62.webp\"\/>Table 1: Recall values for prompt-based log generation.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-63.webp\" alt=\"\" class=\"wp-image-147326 webp-format\"  data-orig-src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/image-63.webp\"\/>Table 2: Recall values for agentic workflow-based log generation.<\/p>\n<p class=\"wp-block-paragraph\">Across the evaluation datasets we used, AI\u2011driven synthetic log generation shows strong potential to produce semantically meaningful logs from TTPs and attacker actions. It can capture multi\u2011event sequences, preserve parent\u2011child process relationships, and generate realistic command lines. <\/p>\n<p class=\"wp-block-paragraph\">This capability can accelerate detection engineering by reducing dependence on costly lab setups and enabling rapid experimentation, without sacrificing realism or safety.\u00a0Our early experiments with reinforcement learning with verifiable rewards also look\u00a0promising and\u00a0could improve verbatim alignment when sufficient training data is available.\u00a0<\/p>\n<p>References<\/p>\n<p>ATLASv2: ATLAS Attack\u00a0Engagements,\u00a0Version 2:\u00a0<a href=\"https:\/\/arxiv.org\/pdf\/2401.01341\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">2401.01341<\/a>\u00a0<\/p>\n<p class=\"wp-block-paragraph\">This research is provided by Microsoft Defender Security Research with contributions from Raghav Batta and\u202f members of Microsoft Threat Intelligence.<\/p>\n<p>Learn more<\/p>\n<p class=\"wp-block-paragraph\">For the latest security research from the Microsoft Threat Intelligence community, check out the\u00a0<a href=\"https:\/\/aka.ms\/threatintelblog\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Microsoft Threat Intelligence Blog<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">To get notified about new publications and to join discussions on social media, follow us on\u00a0<a href=\"https:\/\/www.linkedin.com\/showcase\/microsoft-threat-intelligence\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">LinkedIn<\/a>,\u00a0<a href=\"https:\/\/x.com\/MsftSecIntel\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">X (formerly Twitter)<\/a>, and\u00a0<a href=\"https:\/\/bsky.app\/profile\/threatintel.microsoft.com\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Bluesky<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">To hear stories and insights from the Microsoft Threat Intelligence community about the ever-evolving threat landscape, listen to the\u00a0<a href=\"https:\/\/thecyberwire.com\/podcasts\/microsoft-threat-intelligence\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Microsoft Threat Intelligence podcast<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">Review\u202four\u202fdocumentation\u202fto learn\u202fmore about our real-time protection capabilities and see how\u202fto\u202fenable them within your\u202forganization.\u202f\u202f\u00a0<\/p>\n","protected":false},"excerpt":{"rendered":"In this article Logs and telemetry\u00a0are\u00a0the\u00a0foundation\u00a0of modern cybersecurity. They enable threat detection, incident response, forensic investigation, and\u00a0compliance across&hellip;\n","protected":false},"author":2,"featured_media":439257,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-439256","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/439256","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=439256"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/439256\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/439257"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=439256"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=439256"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=439256"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}