Everything here happened inside bug bounty programs, in scope and with permission. I reviewed every finding and submitted it by hand. Third-party data is redacted, and the examples were disclosed before publication.

For the past few months, I have been testing whether AI can handle the repetitive parts of bug bounty hunting without deciding on its own that a bug is real. Mapping sites, reading JavaScript, comparing requests and revisiting old leads takes time, so I wanted a helper that could keep doing that work and bring me only what deserved a closer look.

That idea became my hackbot, and the first version could explore targets and draft reports. Early on, it reported what looked like a critical IDOR, an access-control bug where one user can reach another user’s data, but when I repeated the request, both notes belonged to the same account. There was no second user and no bug, which gave the project a simple rule: a lead only moves forward when I can reproduce it from saved evidence.

The harness enforces that rule by choosing tools, tracking which account sent each request and checking what came back, and so far it has narrowed 642 leads to 167 that passed the proof check; I submitted 121, and programs accepted 85.

How a lead reaches me

From reports to repeatable proof

What v1 got wrong.

My first version was a prompt wrapped around tools: a discovery pass found websites, APIs and JavaScript files, then the model picked a test, sent a request and wrote a report.

It found real issues, but it also confused clues with proof:

  • HTML containing <script> became “confirmed XSS.” XSS means making attacker-controlled JavaScript run in another user’s browser, but here the browser treated it as harmless text.
  • A publishable browser key became a leaked secret.
  • A 200 response containing a login page became an authentication bypass.
  • A firewall blocking a test became evidence that the test had reached the database.
  • Reading an object created by the same account became cross-user access.

In each case, the model saw a clue and described it as impact, and asking it to check again usually produced the same conclusion.

The oracle.

I stopped asking whether a bug looked real and gave every bug type a pass-or-fail check. I call it an oracle: a specific effect that should happen only when the bug is real.

  1. Check what the application did, not how the model described it.
  2. Compare the test with a normal request, or use a second account or a one-time random value.
  3. Repeat the test and expect the same result.

The model suggests the bug, while ordinary code checks the proof.

what counted as proofone pass-or-fail rule per bug type
Missing login checkCount it: protected data or an action works without logging in and works again.Do not count: a login page with status 200 or an intentionally public route.
Broken permissionsCount it: a limited account completes a restricted action and the change can be read back.Do not count: a hidden button or a different response size.
Exposed sensitive dataCount it: a request without login returns a small, valid protected record.Do not count: filenames, empty storage or public metadata.
Package source mix-upCount it: an approved test package calls back from the internal build.Do not count: finding a free package name by itself.
Login bypassCount it: an invalid login token completes a protected action and the change can be read back.Do not count: a token error or status 200 by itself.
Unsafe configurationCount it: a specific read or safe write works while the normal comparison is denied.Do not count: a public key, banner or version number.
Cross-account accessCount it: test account B reads an object created by test account A.Do not count: one account reading an object it created.
Server-side requestCount it: the target contacts a private test URL containing a fresh random value.Do not count: a repeated URL, a request from my own machine or a timeout.
The eight types the harness checks most often. Anything that does not clear its row stays a lead.

One rule comes before the others: a status code alone is never proof. A 200 says the server returned something. It does not show that access was allowed or that the payload ran.

For timing bugs, the checker alternates the suspicious request with a normal one and ignores unstable servers. For access-control bugs, it records which account created an object and which account read it. For XSS, the JavaScript must actually run in a browser. Each rule works like a small software test:

def prove_cross_account_read(created, replay, normal_request):
    return all([
        created.owner == "account_a",
        replay.reader == "account_b",
        replay.object_id == created.object_id,
        replay.body.owner == "account_a",
        normal_request.status in {401, 403, 404},
    ])

The required proof is saved before the model starts testing:

bug_type: cross_account_read
object_created_by: account_a
object_read_by: account_b
request_fingerprint: 7c1d…
proof:
  owner_in_response: account_a
  normal_request_status: 403
  repeated: 3_times
verdict: proven

Those account fields rejected the fake bug from the opening. The creator and reader were the same account, so the finding stopped before a report existed.

The proof changes with the bug. A server-side request must reach a private test URL, browser code must run, and a crash must happen again with the same input. The evidence comes from the application, not the model’s explanation.

Building a harness that could say no

V1 watched for new websites, API routes, input fields and JavaScript changes. jsluice and Semgrep pulled interesting paths from code, then separate workers tested access control, login flows, race conditions, old files, server-side requests, single sign-on and exposed software.

The coverage was useful, but the shared records were weak. One proof label could point to several reports, two websites with the same path could be treated as one, and a rejected lead could return later as if it had been confirmed. More tests would not fix that, so I gave every lead, request and proof its own record.

The current harness runs as a loop on a small cloud server rather than one large agent. Discovery records where each route came from, then ordinary code turns it into a small question: can account B read account A’s object? Does the page change after login? What can this public key do?

The model sees only that question, not the whole target. A separate checker repeats its request without reading the model’s explanation. It returns rejected, needs-proof or proven; I decide what becomes a report.

Each saved lead carries enough information to survive a restart without relying on the model’s memory:

{
  "lead": "api.example.com|GET|/v2/accounts/{id}|access-control",
  "found_in": "main.8f31.js:18422",
  "owner_account": "session:7d2c",
  "reader_account": "session:b915",
  "test_request": "3f38…",
  "normal_request": "b87a…",
  "proof_needed": "response still names the owner account",
  "state": "needs-proof"
}

Recent catches from the harness cockpit, each row carrying its proof state

A plain-language view of the queue: severity on the left, proof status on the right. I do not report a lead while it is still being checked.

Login state mattered too. On one target, tokens expired after four minutes and could only be renewed in the browser. The agent kept testing while logged out. I now refresh that session every 150 seconds, pass the current token into each test and record which account owns each cookie. If login fails, the run pauses.

What still broke.

Once real findings started arriving, I stopped adding features and audited the harness itself.

False proof rules.

The fake IDOR from the opening became my first saved test. The new check requires two accounts and records who created the object before the other account tries to read it.

A similar mistake affected SSRF, a bug where an application can be tricked into making a server-side request. My checker searched the response for the URL it had just inserted, so a page that merely repeated the URL looked vulnerable. Now proof requires the target to contact a private test URL containing a fresh random value.

Template injection had the same problem. Testers often send 7*7 and look for 49, but my rule accepted any 49 in the response, including IDs and timestamps. It produced 25 false criticals. Large random numbers and a normal comparison request fixed it. The checker also learned to reject firewall block pages.

The worst failure was quieter: two renamed fields made one check always pass. I fixed it by keeping five known false alarms, including a login page, a firewall block and an object read by its owner. After every code change, the harness tests those examples again. If any of them becomes “proven,” I know the new code is broken.

One website had produced 450 leads from fake “not found” pages and version guesses. After filtering them, 22 useful leads became visible. Removing noise made the queue easier to work as well.

The rule I kept from the audit:

If a rule must never change, put it in code. Prompts can drift; a failed software test is obvious.

Where real findings disappeared.

False positives were only half the problem. Generation and validation worked, yet good findings still disappeared before submission.

  • needs-proof was a dead end, so leads that needed one more request were never revisited. Saving the exact request and response for each decision recovered 23 reports.
  • A quality score would have blocked all 42 accepted findings in one audit sample. I stopped using the score as a hard cutoff and used it only to decide what I reviewed first.
  • A “nothing left to test” rule counted firewalls and expired logins as failed tests. Four of five runs stopped early and left 639 leads untouched.
  • The duplicate checker ignored the website name, so login.example.com/api/v1 and admin.example.com/api/v1 became the same route. Fixing that mistake raised measured coverage from 20.3% to 62.7%.

I had counted what entered the pipeline but not why it left. Every step now records a reason: duplicate, outside the program, login expired, comparison failed, proof failed, rejected, submitted or accepted.

One v1 run made the waste easy to see. The harness called model workers 111 times across 37 groups of tasks. It spent $106.70, added five report files and still left 73 leads waiting. Repeating work under vague proof labels cost another $21.91, so I fixed the records before buying more model time.

A museum displaying lines of code, story points, pull requests and tokens spent as meaningless metrics

Activity is easy to count. I care about the smaller number of results that can be repeated and safely reported.

Choosing the models

Model routing means choosing a model for each job instead of using the most expensive one everywhere. V1 reached for the strongest model too often, even when it was only summarizing a JavaScript file or grouping duplicate reports.

one job, one clear owner
Routine explorationSonnetRead JavaScript, summarize routes and try small variations.
Uncertain casesOpusReason through tests with several steps, accounts or conflicting responses.
Harness maintenanceCodexTurn past mistakes into automated tests and patch the surrounding code.
ProofOrdinary codeRepeat the request and check the exact effect required for that bug type.
SubmissionMeReproduce the result, reduce impact, redact data and write the report.
The models help at different stages. None of them gets to approve its own finding.

I keep each task small. Devansh’s needle-in-the-haystack writeup shows why a focused question beats dumping every file into the context. Andy Gill’s harness overview explains the surrounding software, while Ethiack’s benchmark separates planning from execution.

As of September 2026, Anthropic lists Sonnet 5 at $2 per million input tokens and $10 per million output tokens. Opus 5 costs $5 and $25. Input tokens are the text the model reads; output tokens are the text it writes. Both models support up to one million tokens of context, but a large limit is not a reason to fill it. (Anthropic model overview)

model_router:
  normal_hunting: sonnet-5
  use_opus_when:
    - test_has_three_or_more_steps
    - test_crosses_account_boundary
    - responses_disagree
harness_code: codex
proof_check: ordinary_code
submission: human

At the same amount of text, sending 80% of the work to Sonnet and 20% to Opus costs about 48% less than using Opus for everything. That is only a budget estimate. A cheaper model can still cost more if it needs extra attempts or creates more false leads.

Hacktron saw exactly that in its oauth2-proxy experiment. Sonnet 4.6 was cheaper per token, but its full run cost more because it produced more leads and triggered more follow-up work. In the table below, 5/7 means the setup found that known bug in five of seven runs.

Model they testedKnown bug AKnown bug BLeads per runCost per run
Claude Opus 4.67/75/7~260~$79
Claude Sonnet 4.65/62/6~1,120~$122
GPT-5.4 Mini7/106/10~490~$10
Gemini 3.1 Flash9/1010/10~264~$3.70

Those numbers describe their harness and two known bugs, not a universal ranking. The extra leads explain why Sonnet cost more even though its tokens were cheaper: every lead created more work downstream.

Gemini 3.1 Flash has the strongest result in that table. I have not switched because I have not tested it inside my harness. First, I would run it on the same saved tasks, with the same proof budget, and compare how many useful leads survive.

Codex has a different job. I use it to maintain the software around the hunter, not to decide whether a bug is real. When a checker fails, I save the bad example and the expected answer. Codex can write a test, patch the checker and rerun the saved cases. OpenAI describes GPT-5.3-Codex as an agentic coding model with a 400,000-token context window. (official model page)

I have not benchmarked Codex as a hunter. That comparison would need the same target data, proof budget and blind review.

What the numbers show

The first 642 leads were collected over four months, while some program decisions arrived later. Of those leads, 508 were worth a closer look and 167 passed the oracle. By September 2026, I had submitted 121 and programs had accepted 85.

what happened to 642 leads
leads found642
worth a closer look508
oracle-proven167
submitted121
accepted by programs85
Updated September 2026

Most of those reports came from v1, before the memory and oracle fixes. The last 24 leads to pass the oracle came after them, and all 24 were submitted and accepted. That is one small cohort, not a success rate, and I will not know whether the fixes helped until several more months look like it.

Programs can reject a report because it is a duplicate, outside their rules or rated differently. That is why I keep leads, proven bugs, submissions and acceptances separate.

The harness starts with checks it can repeat reliably. I still review unusual business rules and timing bugs myself because I do not have a general automatic proof for them.

The bugs that made it through.

Below are 26 accepted reports from the first four months. Individual rewards are rounded or masked where the program is private.

FindingBug typeWhereSeverityReward
Bulk export of 2,928,956 customer records, effectively the whole databasePII exposureCarmaker, PeruCritical$X,000
963,705 cross-customer payment transaction recordsPII exposurePayments platformCritical$X,000
National civil registry lookups for any citizen through authenticated IDORPII exposureFintualCritical$X,000
More than 65,000 customer records across 22 entities through an unauthenticated CRM endpointPII exposureBank, ChileCritical$X,000
National ID photographs and home addresses available without loginPII exposureBank, GuatemalaCritical$X,000
5,279 landowner records, including national ID numbersPII exposureEnergy company, PeruCritical$X,000
Signed contracts and financial documents through sequential IDORPII exposureDocument platformCritical$X,000
Full names and tax IDs for about 734,000 users through a P2P order endpointPII exposureLemonHigh$X00
PII for about 180,000 merchants available without loginPII exposureFintualHigh$X00
131,374 vehicle and owner records enumerable by sequential IDPII exposureCarmaker, PeruHigh$X00
Customer insurance policies in a public S3 bucket, covering 640,550 objectsPII exposureInsurer, PeruHigh$X00
1.3 million vehicle-owner records through a plate lookup in a booking flowPII exposureCarmaker, PeruHigh$X00
IDOR in GraphQL hirerUser queries allowing mass employer PII enumerationPII exposureSEEKHigh$700
An unclaimed internal npm scope installed by CI and production systemsRemote code executionCoopeuchCritical$X,000
61 internal package names unclaimed across PyPI, RubyGems and npmRemote code executionFintualCritical$X,000
Public backend key with row-level security disabled and write access to the production feedWrite accessLemonCritical$750 + 1,500 USDC
Full card numbers exposed without login through integer type confusionCard-data exposureBank, PeruCritical$X,000
Production database credentials returned in a Set-Cookie headerCredential exposureBank, ChileCritical$X,000
Patient health declarations could be written and deleted without loginData integrityInsurer, ChileCritical$X,000
Outside subscriptions allowed on more than 390 production SNS topics across three AWS regionsBroken permissionsInsurer, PeruCritical$X,000
Real users’ session cookies exposed through a public archiveAccount takeoverMarketplaceCritical$X,000
OAuth wildcard redirect URI enabling account takeoverAccount takeoverHR platformCritical$X,000
Payment API checked for a bearer token but did not verify its signatureLogin bypassProcessor, PeruCritical$X,000
SAP RFC command injection through an exposed SOAP endpointRemote code executionTelecom application, PeruCritical$X,000
SAML configured with WantAssertionsSigned=falseLogin bypassTelecomHigh$X00
Subdomain takeover caused by insufficient validationSubdomain takeoverFoxy.ioHigh$500

Together, the rewards in this table added up to just under $30,000 during the four-month test. The redacted cells will not visibly add to that total. I include it because it answers a practical question: the harness produced findings that programs considered useful. It is not a stable monthly rate, and it does not show which model deserves credit.

Most of the value came from recording four ordinary facts: who sent the request, where an identifier came from, what changed compared with the normal request, and whether the result happened again.

Three findings that changed the harness

From open signup to account takeover

An investment platform used Amazon Cognito to manage logins for an internal Kibana dashboard. Cognito handled the accounts; Kibana displayed application and email logs. The signup form accepted any email address instead of limiting access to company staff.

The first clue was only “internal dashboard exposed.” To show the real impact safely, I requested a password reset for my own account and watched the new reset link appear in the email log. That proved the dashboard exposed live messages, not only an old archive, without touching anyone else’s account.

The harness now treats old records and newly triggered messages as different levels of proof.

When a private build installed a public package

A contractor’s public repository exposed an .npmrc, including Coopeuch’s private package server and credentials. Within the program’s rules, those credentials let me confirm the exact internal package and version in the private panel.

The public registry returned the same 400 response for real and invented names, so I could not use it as proof that a package was free. After the program approved a controlled test, I published a harmless package with a unique callback. Six internal build hosts installed it within fourteen minutes; four ran it as root.

I changed the harness to confirm private packages through an authenticated source, discard generic registry errors and require an internal installation callback as proof.

Full chain in The Version Number That Got Me Root Inside a Bank.

A tweet that reached eToro’s account tool

eToro’s AI assistant could read public pages and use tools as the logged-in user. A tweet contained instructions that made the assistant call an account tool and try to place its output in an external URL. This is prompt injection: the model treated public text as a command.

I fixed this by separating content from permission: public text may be read, but only the user can authorize a private tool call. The harness records where each instruction came from and blocks public content from choosing a private tool or an external destination.

Detailed writeup: Hacking eToro’s AI Assistant.

What I’d do differently

If I started again, I would keep the first version small:

  1. Start with three bug types that have clear proof.
  2. Save every request, response and account identity.
  3. Compare each suspicious request with a normal one.
  4. Keep the model that finds a lead separate from the code that checks it.
  5. Reproduce and submit every report myself.

I would also avoid giving the model a large library of old reports. It started copying familiar bug patterns instead of paying attention to the application in front of it. Short instructions written for the current test worked better.

Closing thoughts

The useful part of an LLM workflow is rarely the model alone, but the structure around it: how work is split into stages, what context reaches each stage, which tools are available, how results are challenged and what gets remembered for the next run. A good harness does not replace judgment or validation; it makes both repeatable, which is why the model explores, code checks the evidence and I decide what gets reported.

That separation mattered more than changing models because it also changed what I measure. A large queue of leads may look impressive, but a smaller set of findings I can replay and defend is more useful; for this project, the proof process became the product.

The harness is private for now, although this post shares its design and the mistakes behind it.

By the way, I see many people build Claude skills with a vague description and far too many instructions. A useful SKILL.md can still be short:

---
name: triaging-bug-reports
description: Validates authorized web-security findings with repeatable evidence. Use when triaging a suspected bug or preparing a bug-bounty report.
---

# Triage a bug report

Treat the model's finding as a lead, not proof.

1. Confirm that the target and test are authorized, then load the original request, response and tool output.
2. Record which account sent the request and which account owns the object.
3. State what must happen for the bug to be real, then choose a check that ordinary code can evaluate.
4. Replay the test and send a normal comparison request. Repeat both if timing or unstable responses could affect the result.
5. Use a bug-specific check: a second account for access control, browser execution for XSS, or a unique callback for server-side requests.
6. Reject status codes, reflected text, login pages and firewall blocks when they do not prove impact.
7. Save the requests, responses, identities and timestamps needed to reproduce the result.

Return `proven`, `needs-proof` or `rejected`, followed by the reason, evidence and missing proof. Draft a report only when the result is `proven`.

The important part is the description: say what the skill does and when Claude should load it. Save the file as .claude/skills/triaging-bug-reports/SKILL.md. (Anthropic’s skill-authoring guide)

Until the next one, stay curious, stay ethical.