SecurityClaw: Autonomous Bug Bounty Hunting Platform
SecurityClaw is an autonomous penetration testing platform built from the ground up for bug bounty hunting on Intigriti and HackerOne. It orchestrates 55+ specialised security skills — from passive reconnaissance through active exploitation — coordinated by an AI layer that generates hypotheses, validates findings, replans mid-campaign, and learns from every outcome.
This isn't a smarter wrapper around Nessus. It's a modular pipeline where each skill is purpose-built for a specific attack class, real industry-standard tools (nmap, nuclei, ffuf, sqlmap, Metasploit) do the execution, and an AI layer makes the decisions: what to run next, which findings are real, which chains are worth pursuing.
This page covers the current architecture, the AI pipeline (Phases A–D), the skills inventory, and how campaigns actually run.
What SecurityClaw Is
SecurityClaw is a Python-based platform with a modular skill architecture. A "skill" in SecurityClaw is a self-contained module — a Python class with a SKILL_MANIFEST declaration and an execute() method. Skills range from passive OSINT (certificate transparency monitoring, subdomain enumeration) to active exploitation (SQL injection, IDOR, SSRF chaining) to cloud-specific scanning (AWS bucket enumeration, Azure naming attacks).
The differentiator vs a tool runner: skills don't execute in a fixed sequence. The adaptive campaign engine maintains a dynamic queue. Each skill's output produces typed findings — SUBDOMAIN_FOUND, OPEN_PORT, CREDENTIAL_FOUND, WAF_DETECTED — and chain rules fire automatically when findings match trigger conditions. A new subdomain discovered mid-campaign immediately triggers nuclei scanning on that subdomain. A detected WAF triggers a WAF-aware scan strategy. An internal hostname leak triggers the SSRF hypothesis chain.
The campaign never just follows the script. It follows the evidence.
EC2 Campaign Execution — Ephemeral Infrastructure Per Campaign
SecurityClaw runs campaigns on ephemeral AWS EC2 instances. This wasn't a design convenience — it was a technical requirement. Real bug bounty targets block residential IPs, flag repeated requests from a known range, and employ WAFs that profile scanning behaviour over time. A platform running from a fixed IP hits those walls immediately.
The solution: spin up a fresh EC2 instance (Amazon Linux 2023, t3.medium, eu-west-1) for each campaign, run the full campaign on it, capture results, and terminate. The instance — and its IP — exists only for the duration of the campaign.
The full EC2 campaign lifecycle runs through ten phases:
| Phase | What Happens |
|---|---|
| PLAN | CampaignDirector validates the campaign plan JSON. Human approval gate for exploitation-class skills. |
| UPLOAD | Full SecurityClaw repo tarballed and uploaded to S3. Campaign plan uploaded separately. |
| LAUNCH | EC2 instance launched with IAM profile, security group, and UserData bootstrap script. |
| WAIT_READY | Polls until instance is running with healthy system status checks. |
| SSM_WAIT | Waits for AWS Systems Manager agent to come online on the instance (replaces SSH). |
| BOOTSTRAP | UserData installs dnf packages, Go tools (subfinder, httpx, nuclei, ffuf), Python 3.11. |
| EXECUTE | SSM sends the campaign command. Output streamed to S3 and CloudWatch Logs. |
| CAPTURE | Results read from S3 (avoids 24K truncation). Findings and skill results parsed from delimited JSON blocks. |
| TERMINATE | Three-layer termination guarantee: Python finally block, UserData shutdown timer, EventBridge reaper. |
| REPORT | Results written locally. AI attack chain narrative generated. |
A key architectural choice: command execution runs over AWS Systems Manager (SSM), not SSH. SSM is session-resilient — if the orchestrator process dies mid-campaign, the campaign continues running on the EC2 instance. The idempotent state machine execution driver (campaign_runner.py) picks up from the last persisted phase when the orchestrator restarts.
Cost per campaign: approximately $0.05/hr on a t3.medium. A typical recon-to-exploitation campaign runs 60–90 minutes. Each campaign gets a fresh public IP, full isolation from other campaigns, and no history that a target's WAF can correlate.
The AI Pipeline — Four Phases of Machine Reasoning
The platform's AI layer is organised across four phases, each adding a distinct layer of reasoning over raw tool output.
Phase A — Bedrock-Powered Finding Intelligence
FindingValidator runs after every skill execution, before findings are committed to the database. It calls a Bedrock Haiku-class model to classify each finding:
CONFIRMED— high confidence real findingLIKELY_VALID— probable, needs manual verificationLIKELY_FP— probably a false positiveFALSE_POSITIVE— discard
Every finding that enters the database carries validation_status, validation_reasoning, and validation_confidence. This is what keeps a high-volume campaign from drowning the analyst in noise.
AttackChainAnalyser runs post-campaign using a Bedrock Sonnet-class model. It takes the full validated finding set and identifies multi-step attack chains — e.g. "SSRF via leaked internal hostname → internal API access → credential exfiltration" — and generates human-readable narratives. Results are stored in the campaign_chains table and surface in the final report.
Phase B — Adaptive Execution Engine
The AdaptiveCampaignEngine is the core execution loop. It maintains a queue of SkillTask objects and processes them with safety limits: max chain depth 5, max 50 skills per campaign, deduplication by (skill_id, target), 2-hour timeout. When a skill produces a typed finding that matches a chain rule, the follow-up skill is automatically queued.
Built-in chain rules include:
OPEN_PORT→redis-checkFILE_FOUND→file-content-analyserCREDENTIAL_FOUND→hashcat-crackSUBDOMAIN_FOUND→nuclei-scanWAF_DETECTED→cf-aware-scan-strategyInternalHostnameLeak→ SSRF chain
Other Phase B capabilities: JSBundleAnalyzer uses tiktoken-chunked Bedrock calls to extract hardcoded API keys, client IDs, and internal endpoints from JavaScript bundles. The AI IDOR Scanner runs three-phase detection — resource enumeration, cross-user access testing, and Bedrock differential analysis — to confirm IDOR vulnerabilities with high precision.
Phase C — Autonomous Intelligence Loop
Phase C is where SecurityClaw stops being a sophisticated scanner and starts behaving like an autonomous researcher. Ten components work together:
C1 — AdversarialHypothesisEngine: Before running any tools, the platform forms named attack hypotheses with confidence scores. "H1: Broken Object-Level Authorization on /api/users endpoint — confidence 0.7." These hypotheses guide which evidence to collect first and produce resolution narratives after the campaign completes.
C2 — RePlanner: Checked after every skill batch. Seven trigger conditions — high-severity finding, new subdomain, auth bypass confirmed, SSRF confirmed, API keys found, mobile app detected, internal service fingerprinted. AI decides: CONTINUE / PIVOT / PRUNE / STOP. PIVOT adds new skills. PRUNE removes low-value queued skills. Exploitation skills cannot be auto-approved via replanning — always require human sign-off.
C3 — IntelligenceStore: SQLite database recording completed campaigns, finding patterns, tech tags, and skill hit rates. Feeds C1 hypothesis confidence and C7 program selection. Compounds across every campaign — the platform gets smarter with each run.
C4 — SkillGap Detection: Automatically detects three gap types — attack classes with no skill coverage, findings that can't be exploited, and blind spots where skills with high hit rates aren't running. Gaps auto-close when C5 discovers a covering skill.
C5 — SkillRegistry: Auto-discovers skills at startup by scanning for SKILL_MANIFEST declarations. No imports at discovery time. Syncs gap state with C3.
C6 — OutcomeSyncJob: Polls Intigriti API for submission outcomes. Accepted findings increase C3 pattern weights; rejections decrease them. Closes the real-world feedback loop into the intelligence store.
C7 — ProgramSelector: Scores Intigriti programs against C3 historical hit rates. Identifies which programs SecurityClaw's skill set is best positioned to find findings on. Weekly cron with AI-generated rationale.
C8 — SpecGenerator: Auto-generates structured skill specs for C4-detected gaps. Outputs to docs/specs/auto-generated/. Creates a documented backlog of skills that need writing.
C9 — SubmissionDrafter: Generates Intigriti-formatted bug reports for MEDIUM+ validated findings. requires_review: True always — no auto-submission without human sign-off.
C10 — SkillHealthMonitor: Tracks per-skill health: run count, finding rate, consecutive empty runs. Detects hard failures, soft degradation, and stale skills. Posts nightly dashboard to Slack.
Phase D — Production Gap Skills
Ten class-based skills that fire automatically to cover capability gaps identified during live campaign testing:
| Skill | Gap Addressed | Chain Rule Trigger |
|---|---|---|
| HeaderAnalysisSkill | Security header coverage | SUBDOMAIN_FOUND |
| ProgramScopeManager | Multi-TLD scope setup | Pre-campaign |
| ToolAvailabilityMonitor | Pre-flight binary check | Pre-campaign |
| WAFDetectionSkill | WAF identification | Early campaign |
| SensitiveServiceFingerprintSkill | Internal service detection | SUBDOMAIN_FOUND |
| APIKeyValidatorSkill | Key validation post-discovery | JS_BUNDLE_FOUND |
| MobileAppAnalysisSkill | Mobile attack surface | MOBILE_APP_DETECTED |
| SSRFHypothesisSkill | SSRF chain on internal leaks | InternalHostnameLeak |
| IntigritiScopePuller | Live scope from Intigriti API | Pre-campaign |
| CrossAssetCorrelator | Cross-subdomain correlation | SUBDOMAIN_FOUND |
Skills Inventory — 55+ Specialised Modules
SecurityClaw's skill library covers the full web and cloud attack surface. Skills are divided into class-based (Python, auto-discovered via SKILL_MANIFEST) and legacy subprocess skills (external tool wrappers registered in the executor). All class-based skills have dry-run mode and are tested without live network dependencies.
Class-Based Skills (auto-discovered)
- Recon & OSINT: certificate-transparency-monitor, cross-asset-correlator, intigriti-scope-puller, js-bundle-analyzer, js-bundle-recon, js_analyzer, nextjs-recon, sensitive-service-fingerprint, tech-stack-cve-scanner, webmail-cve-fingerprint, whitelabel_detector
- Web Application: auth-portal-tester, business-logic-scanner, header-analysis, idor-scanner, logging-monitor-checker, open-redirect-probe, open_redirect_hunter, security-header-checker, security-misconfiguration-checker, session-security-tester, sqli-probe, sri-checker, ssrf-hypothesis, ssrf-probe, tls-crypto-auditor, waf-detection, xss-probe
- API Security: api-key-validator, api-schema-discovery, apigw-cors-tester, authenticated_api_sweep, oauth-client-enum, oauth-security-analysis, oauth_scope_abuse, swagger_admin_probe
- Cloud: appinsights_telemetry_probe, azure_naming_enum
- Mobile: mobile-app-analysis
- Infrastructure: ec2_campaign_executor, program-scorer, hunt_orchestrator
Legacy Subprocess Skills (tool wrappers)
- Recon: nmap-recon, masscan-fast, shodan-intel
- Enumeration: gobuster-enum, ffuf-fuzz, enum4linux-smb
- Scanning: nikto-scan, nuclei-scan
- Exploitation: sqlmap-injection, hydra-bruteforce, hashcat-crack, metasploit-exploit
- Post-Exploitation: proxychains-stealth, sliver-c2, aircrack-wireless
Exploitation-class skills (metasploit-exploit, sliver-c2, sqlmap-injection, hydra-bruteforce, hashcat-crack) require requires_approval: true in the campaign plan. They will not execute without explicit human sign-off — this is a hard gate, not a configuration option.
Quality and Test Coverage
SecurityClaw maintains a 98%+ test coverage gate on the skills/ directory. The test strategy:
- No live network dependencies — all tests use mocked requests, mocked subprocess calls, and mocked Bedrock responses. The full test suite runs offline.
- Parametrised dry-run tests — every class-based skill is covered by a parametrised
test_dry_run_*.pytest confirming it returns an empty list without making network calls whendry_run=True. - Phase C component tests — dedicated test files with heavy mocking of IntelligenceStore, Bedrock clients, and Intigriti API responses.
- CI pipeline — isort + black + flake8 + mypy strict + radon complexity + bandit + pip-audit + pytest. Every PR must pass all gates before review.
How a Campaign Actually Runs — End to End
A full SecurityClaw campaign against a bug bounty target follows this sequence:
- Pre-campaign setup: IntigritiScopePuller fetches the live program scope and Rules of Engagement from the Intigriti API. ProgramScopeManager validates multi-TLD targets. ToolAvailabilityMonitor confirms all required binaries are present on the execution instance.
- Hypothesis generation: AdversarialHypothesisEngine generates 3–7 named attack hypotheses based on the target, prior campaign history from IntelligenceStore, and common patterns for the target's tech stack.
- Execution loop: AdaptiveCampaignEngine processes the skill queue. Each skill result is validated by FindingValidator, chain rules fire for matching typed findings, and the RePlanner evaluates its 7 trigger conditions after each batch.
- Post-campaign analysis: AttackChainAnalyser generates multi-step attack narratives. IntelligenceStore records the campaign. Gap detection runs. SpecGenerator documents any new coverage gaps.
- Submission: SubmissionDrafter generates Intigriti-formatted reports for all MEDIUM+ confirmed findings. Human reviews each draft before submission. OutcomeSyncJob feeds outcomes back into the intelligence store.
SecurityClaw vs Vulnerability Scanners
The distinction matters because the bug bounty programmes SecurityClaw targets have already been scanned by automated tools. Finding something worth a payout means finding something those tools missed.
| Vulnerability Scanner | SecurityClaw | |
|---|---|---|
| Execution | Fixed scan sequence | Dynamic queue, chain-rule driven |
| IP exposure | Fixed source IP | Fresh EC2 IP per campaign |
| Finding quality | Raw output, high noise | Bedrock-validated per finding |
| Finding chaining | None | Typed findings + chain rules |
| Mid-campaign adaptation | None | RePlanner (7 trigger conditions) |
| Learning across campaigns | None | IntelligenceStore compounding |
| Attack narratives | CVE list | Multi-step chain analysis (Sonnet) |
| Session resilience | N/A | SSM + idempotent state machine |
Current Status and Known Gaps
As of April 2026, SecurityClaw is operationally deployed and running campaigns. Key items in active development:
- Bootstrap tool coverage: The current EC2 UserData installs nmap, subfinder, httpx, nuclei, and ffuf. A second wave of bootstrap tools (nikto, gobuster, sqlmap, hydra, hashcat, masscan, cloud scanning tools) is in development to bring the full subprocess skill set online in EC2 campaigns.
- Intelligence store population: Phase C's AI features (hypothesis confidence, program scoring, replan heuristics) improve with each completed campaign. Initial campaigns are actively building the history that makes these features increasingly accurate.
- Plan generation: Campaign plan generation (target domain → structured campaign plan JSON) is handled by the security engineer prior to launch. An automated planning layer is on the roadmap.
For the full technical deep-dive and gap analysis, see the security research documentation.
Live Campaign Stats
| Category | Campaigns | Pass | Partial | Fail | Tools |
|---|---|---|---|---|---|
| Web Scanning | 1 | 100% | 0% | 0% | multi_tool_campaign |
What This Means for Bug Bounty
Bug bounty programmes have been live long enough that the easy wins — exposed admin panels, default credentials, obvious SQL injection endpoints — are gone. What remains requires chaining: a subdomain found via certificate transparency that runs an old CMS, a JWT misconfiguration discovered in a JS bundle, a leaked internal hostname that enables SSRF to an internal metadata service.
That's exactly what SecurityClaw is built to find. The chain rules, the typed finding system, the adaptive replanning — all of it is designed for the multi-step paths that traditional scanners can't see because they don't chain.
The intelligence store means that knowledge compounds. A pattern found on one Intigriti programme — a particular tech stack, a class of vulnerability, a chain that worked — raises the hypothesis confidence for every future campaign that touches a similar target.
This is what an autonomous bug bounty researcher looks like in 2026.