Cold open
417 of 418
Day two. One of my agent lanes is about to report back on a password spray: 418 candidates across 14 addresses, no hits. A clean negative. Cross it off.
Except a check I’d made mandatory fires mid-run: the one credential known to work has started failing too.
At ~28 requests a second, the Burp proxy was answering for the target — its own 1,344-byte “No response received” page, with HTTP 200. 417 of 418 attempts never reached the application. It was reached exactly once.
The agent wasn’t lying. It just couldn’t tell a result from a broken instrument. This is the story of a CTF — and of building the machinery that keeps a team of AI agents honest.
Act 1 · Aug 7–8
The handout
The target: Xenoptic, a fictional frontier-AI company built by the Bug Bounty Village organizers to be broken. Models called Glimpse, Gaze and Augur. A chat assistant, a public API, a developer CLI, two enterprise tenants.
The handout arrived as an AES-encrypted zip that stock unzip couldn’t open, with a brief
and a “binary gift”: 1010 0111. (It turned out to be the first byte of a token baked into
every sandbox.)
About a dozen hosts — chat, api, admin, careers, git, mail, payments, two tenants — all on
one IP, routed by Host header alone. DNS was a wildcard, so the first rule written down
was: resolution proves nothing. Only an HTTP response does.
Hard limits: don’t touch real AWS, delete nothing, no DNS brute-forcing, no social engineering. Everything else was fair game.
Act 2
First blood
The first flag was sitting in a headline field of an unauthenticated staff directory,
/v1/people. Not the most dramatic flag — but that directory became the target list for
everything that followed.
The mail host let anyone anonymously claim any seeded address nobody had claimed yet. A 302-vs-409 response told me which were free.
Password reset trusted delivery to an @xenoptic.ai mailbox — which I now owned.
Request reset, read the token, set a new password. Eight accounts across three
organizations, three of them owners, with zero victim interaction. Meanwhile a disassembled
CLI binary coughed up a hardcoded signing key and the best flag of the event:
flag{d0nt_h4rdc0d3_s1gn1ng_k3ys_1n_th3_cl13nt_s1d3}.
Act 3
Through the AI, into the cloud
The assistant had a web_fetch tool, and a well-built SSRF validator in front of it. Ask it
to fetch a loopback address and it refuses.
But the validator only checked arguments the user supplied. Users could also publish
skills — instruction files the model follows. A URL written into a SKILL.md and emitted
by the model was never checked.
A controlled pair, same account, model, tool and URL, three minutes apart: direct, blocked; through a skill, the live metadata service. The validator’s logic was fine. Its location was the bug.
The credentials it handed over had unrestricted read and write on the entire object store — every tenant, no bucket policy, no access logging. Listing the bucket gave me the key of an executive’s private board deck. Nothing guessed.
Act 4
Apply for a job, own the control plane
Submitting a job application got you a live coding assessment: a Python sandbox that runs your code and returns stdout. No human in the loop.
The internal cloud control plane was protected by where you stood on the network, not by a credential — from the internet it returned a byte-identical 404 with or without the token. The careers sandbox stood inside that boundary, with the hostname and token in its own environment variables.
From inside: a flag in a service banner, a signing secret in CloudWatch Logs, then assume an analytics role and scan a DynamoDB table. Three flags from one job application.
The one control guarding privileged roles was a one-entry deny-list. The role name was
parsed twice: the deny check normalised a full ARN; the permission lookup also accepted a
bare name. RoleArn=xenoptic-secrets-reader passed one and selected the role in the other —
and the response said UnknownRole, so the audit log never recorded it.
Interlude
The map
Put together, it looks like this. Every path starts in the same place: an anonymous attacker on the internet with, at most, a self-registered account. Hover any node to trace its path.
Identity. The open staff directory named the targets. The claimable mailbox and a trusting reset flow took them. An owner’s seat opened an old skill version that still held a secret rotated out of the current one.
Through the AI. Publish a skill, let the model fetch what I wasn’t allowed to, and walk the metadata service to a workload credential that opened the entire bucket.
The paths converge. The control plane wanted two things: a position inside its network, and real credentials. The careers sandbox supplied the first; the SSRF chain supplied the second. Then the parser differential took the one role it was built to refuse.
Git. One leaked profile and one weak password bought push access — enough to prompt-inject the Review Bot, and to merge a private repository into a readable fork.
Fourteen flags sit at the ends of these paths, and every one is a store, not a mechanism: a conversation, a log, a table, a bucket, a bio. The one critical with no store behind it — the one-click device takeover — scored nothing.
Act 6 · Aug 9
No flag, no points
Midway through day two came an operator ruling: the platform doesn’t accept a bug without a flag.
Ten real findings dropped to zero. Among them, an end-to-end supply-chain compromise of the company’s CLI — the most severe thing I found — and forging tool results so the model outranked its own system prompt and handed it over in full.
Looking back, every one of the 14 flags was data at rest that a credential or network position let me read. None came from a mechanism. The question that scores isn’t “what can this credential do?” — it’s “which store does it open, and have I grepped every page?”
Act 7
The machine behind the machine
I ran the engagement with a team of AI agents. It changed shape three times. First, one investigative thread at a time — 48 of them.
Then a panel of five lenses, each an epistemic role rather than a surface. The most valuable was the Falsifier, whose only job was to attack our own conclusions. It overturned three negatives we had banked as fact.
Finally, lanes: a coordinator and parallel agents each owning a surface, with explicit stop conditions and hand-offs. The coordinator’s own briefs were wrong at least six times — and each was caught because lanes re-measured instead of trusting them.
The fix for confident wrong answers was procedural: every negative carries a positive control (could this probe have seen a hit?) and a fabricated control (does invented input get the same bytes?). Byte-identical across real and invented input means a gate, not a result. That trap fired at least five times.
Act 8
Eight deployments
The environment was rebuilt from the same image eight times, each under a fresh 16-character ID. Seeded data and baked-in tokens survived every rebuild; everything I created died with it.
So every session started from a runbook: re-establish capabilities in dependency order. Recon carried over; credentials didn’t.
Over roughly 48 hours: 69 recon threads, 56 findings, 13 formal reports, and 14 flags.
Coda
The playbook
27th to 19th on the closing day. Fourteen flags, nine critical findings.
The rules I’d carry into the next one: on /login, a 200 is an error, not a result. Grep every
response for the proxy’s own error page. Every negative gets two controls. Re-assert a
known-good oracle every few hundred requests. Never falsify an old hypothesis with a later
observation that has its own cause. Hunt stored content, not primitives.
And keep a Falsifier. Parallel agents cover ground fast — but the real work was building the machinery that stops them from reporting confident, wrong answers.