For years, the security community has operated under the mantra that continuous fuzzing is a cornerstone of robust software development. Yet, as many maintainers of large-scale C and C++ projects have discovered, simply enrolling a repository in a service like OSS-Fuzz is rarely a "silver bullet." Even the most well-monitored projects can harbor critical vulnerabilities for years. The persistent bottleneck, according to security researchers, is not the lack of compute power or advanced fuzzing engines, but the intensive manual labor required to maintain them. Developers must constantly monitor coverage metrics, write and refine complex harnesses for hard-to-reach code paths, and painstakingly triage the deluge of crashes generated by modern fuzzers. In essence, effective fuzzing has long demanded a human in the loop.
That dynamic is beginning to shift. Antonio Morales, a researcher at the GitHub Security Lab, has introduced a new autonomous fuzzing framework designed to offload these repetitive, human-intensive tasks to an artificial intelligence agent. The project, titled Fuzzing Taskflow, leverages the GitHub Security Lab Taskflow Agent framework to create a self-contained, end-to-end pipeline. By simply pointing the system at a GitHub repository, the agent takes over the entire lifecycle of a fuzzing campaign: it analyzes the build system, identifies ideal entry points for testing, writes and maintains the necessary harnesses, executes the AFL++ fuzzer, monitors coverage reports, and ultimately triages individual crashes to produce detailed vulnerability reports.
The Fuzzing Taskflow is built on a modular architecture that draws a firm line between decision-making and execution. The LLM agent acts as the "brain," determining which areas of the codebase to target and how to adapt to changing coverage data. Meanwhile, the execution is handled by a set of Model Context Protocol (MCP) tools that provide the primitive building blocks—such as compiling a harness or initiating a fuzzing run. Crucially, the agent does not interact with the low-level tools directly; it orchestrates the pipeline through these predefined primitives, ensuring that the state of the entire operation is managed through a central SQLite database. This design prevents the "memory leaks" of information between stages and ensures that the system can be paused, inspected, and resumed without losing the progress of a campaign.

One of the most significant innovations in this framework is its sophisticated approach to the coverage-feedback loop. In a manual workflow, a researcher typically runs a fuzzer, examines an LCOV report to find uncovered branches, and then manually crafts a new input or harness to push the fuzzer further. The Fuzzing Taskflow automates this iterative dance. By using a "time budget" strategy that scales with the complexity of the task, the system starts with short, high-efficiency bursts and gradually increases the duration for difficult-to-reach code paths. To prevent the system from wasting computational resources on diminishing returns, the framework employs plateau detection. If two consecutive iterations fail to produce a significant increase in line coverage, the agent recognizes that it has hit a limit and pivots to a new strategy or moves to a different target, ensuring the campaign remains efficient.
A common challenge in fuzzing C and C++ is the difficulty of achieving "structure-aware" testing. While basic fuzzers excel at bit-level mutations, they often struggle with complex, text-based formats like XML or JSON. The Fuzzing Taskflow addresses this by integrating four distinct mechanisms for producing structured inputs. First, it utilizes pre-built, format-specific dictionaries and custom mutators for common file types. Second, it performs a static analysis of the target project’s source code, extracting string literals and numeric constants from definitions and enumerations to create a custom, on-the-fly mutator. Third, it employs a dynamic dictionary that grows as the fuzzer encounters new code, enriching itself with tokens found in nearby comparisons and conditionals. Finally, a corpus-splice operator allows the system to recombine existing, high-value inputs—a technique that provides a significant boost over standard bit-flipping mutations.
Maintaining the momentum of a fuzzing campaign is another area where the new framework distinguishes itself. A major inefficiency in many fuzzing setups is the tendency to start from scratch, effectively forcing the fuzzer to rediscover paths it had already navigated in previous runs. The Fuzzing Taskflow maintains a persistent, stable corpus directory for every harness. This allows the system to carry forward successful inputs from past campaigns, ensuring that the work done today compounds with the progress made last week. If a user interrupts a campaign, the system resumes with a "warmed" state, preserving the coverage gains and avoiding the repetitive "cold start" problem.

Perhaps the most daunting task for any security researcher is the triage process. After a fuzzer finds a crash, the developer must determine if it represents a genuine security vulnerability or a benign harness error. The Fuzzing Taskflow automates this by running a multi-stage post-processing sequence. Crashes are automatically minimized using afl-tmin, replayed under AddressSanitizer (ASan) to capture stack traces, and deduplicated using a normalization process that ignores minor differences in stack frames. The agent then analyzes the call chain and the surrounding code, assigning a verdict to each crash.
These verdicts range from confirmed vulnerabilities to harness-induced bugs, and each report includes a detailed root-cause analysis. The agent even suggests a potential fix in the form of a unified diff and a regression test. While developers are strongly cautioned that these AI-generated suggestions are intended as starting points requiring human verification—the agent’s understanding of the code is, after all, limited by the model’s capabilities—the framework provides a remarkably well-prepared foundation for human review.
To ensure transparency throughout this process, the pipeline includes a live HTML dashboard. Accessible on a local port, this interface provides real-time visualization of the campaign, showing coverage progress, the number of crashes discovered, and the current state of the fuzzer. This level of visibility is essential for users who want to monitor the "judgment" of the agent in real time.

The Fuzzing Taskflow is currently configured to use Claude Sonnet 3.5 by default, a choice based on rigorous internal testing, though users have the flexibility to swap models through the configuration files. As the security landscape continues to evolve, tools that bridge the gap between autonomous AI capabilities and the nuanced requirements of software security are becoming increasingly vital. By reducing the reliance on manual labor, the Fuzzing Taskflow aims to democratize the practice of continuous fuzzing, allowing maintainers of projects both large and small to identify and remediate vulnerabilities more effectively than ever before.
For those interested in exploring the framework, the project is open-source and available via the GitHub Security Lab. As the maintainers continue to refine the system, they encourage the community to contribute to its development, test it on their own repositories, and report any issues, further strengthening the collaborative effort to secure the open-source ecosystem.

