By Casey Ellis (Founder of Bugcrowd and The disclose.io Project) in conversation with Anastasiya Kazakova (Cyber Diplomacy Knowledge Fellow and Geneva Dialogue Project Coordinator, DiploFoundation).

Thanks to established good practices, vulnerability discovery has run on a predictable clock: a researcher finds a flaw, a vendor gets time to patch it, and coordinated disclosure (hopefully) manages the interval between the two. This is the most aspirational scenario, one that hasn’t always been followed in reality (unfortunately). With AI, the industry now faces a serious dilemma: AI can now find vulnerabilities almost instantly, while fixing them still takes time for humans. The systems most exposed to that gap, including critical infrastructure, are often the least able to close it.

The Geneva Dialogue on Responsible Behaviour in Cyberspace focuses on stress-testing existing cyber norms and cybersecurity practices against real-world challenges, including emerging technological developments. In this interview, Casey Ellis examines what happens to concepts such as coordinated disclosure, defender-offender parity, and responsibility for the states and private sector once AI systems start finding –and in some cases choosing –vulnerabilities faster than the humans downstream can respond, and argues that the real gap being exposed isn’t primarily technical at all.

The gap is the exposure

If AI can find exploitable vulnerabilities in CI systems far faster than operators can patch them, does that widening gap itself become a new category of cyber harm, even before any incident occurs?

Casey answers plainly: the gap is the exposure. Report volume gets counted because it’s countable, but he argues it’s the wrong number to be watching.

What actually changed, he says, is the clock. Discovery has compressed toward instant while remediation still runs at human speed, and the loop is now too small for most defenders to fit inside. For a critical infrastructure operator, that isn’t a risk of harm – it’s a standing condition of harm. Roughly 29% of known-exploited vulnerabilities were weaponised on, or before the day their CVE published, meaning the patch window closed before it opened. Worse, he notes, the patch itself arms the attacker: fixes don’t reach production at the same speed everywhere, so they reach attackers first and laggards last.

He proposes a name for this condition: remediation deficit, i.e. the measurable distance between when an asset becomes exploitable and when its owner is actually able to act. Measure that, he says, and it becomes governable.

Where coordinated disclosure breaks

Asked whether a case involving an AI system acting without direct human targeting fundamentally breaks the logic of coordinated disclosure, Casey opened with a correction. On the account he read at the time, the cyber refusals had been deliberately switched off for the evaluation, and the model’s goal was to cheat a benchmark rather than attack anyone – a human set the objective, the machine picked the target. That distinction, he argues, is the whole ballgame, and most of the coverage dropped it.

The part that genuinely breaks, he says, sits underneath that distinction. Coordinated disclosure works – when it works – because the researcher chooses it, and the researcher is the only party who can actually be punished. The system runs on a human carrying personal legal and reputational risk – an agent carries none of it. Safe harbour, coordinated vulnerability disclosure (CVD), authorisation: all of it rests on identifying a human who was, or wasn’t authorised. If the actor is a process pursuing a goal nobody wrote down, he notes, it’s no longer clear who the researcher is, or who a vendor would even negotiate with.

He frames this as fixable rather than fatal – though the fix he has in mind addresses the receiving side of the problem, not the deeper question of who counts as ‘the researcher’ when the finder is a process rather than a person.His proposition, built on  disclose.io’s free, copy-pasteable terms, is to create a separate, clearly labeled, intake channel for reports that come from AI agents rather than humans, with a specific, named person responsible for reviewing and acting on them.

In practice, that means a company knows where an AI-submitted report should land and who owns it internally, instead of landing it in the same queue as human-submitted reports, and creating confusion about who sent it and whether it’s worth attention and legitimate. What it doesn’t do, by his own admission, is answer the harder question raised above: who counts as ‘the researcher’ – in a legal, accountability sense – when the finder itself isn’t a person.

A structural asymmetry, not a one-off

Is the asymmetry between guardrailed defensive models and unrestricted offensive ones a one-off problem, or something more structural? Structural, he says, and the asymmetry runs deeper than model policy. Refusals are a product-safety control, not a security boundary: ‘off’ is the adversary’s default condition, open-weight models ship that way, and anyone holding the weights is a fine-tuning away from removing them. A capability that only manifests once refusals are disabled is still a capability, because the adversary was never running your safety stack to begin with, and unrestricted models were going for around EUR 60 a month two years ago.

The deeper asymmetry isn’t the model at all, he argues, but it’s governance. An organisation cannot deploy an autonomous agent into a Fortune 500 network to aggressively rewrite firewall rules in real time. Deploy that same agent as an attacker into the same network, though, and nobody reviews the change control.

So the technology is symmetric – both sides can use equally capable AI agents. But the institutional friction around using it isn’t. Defenders are slowed down by exactly the safeguards (change control, governance, accountability) that make an organisation function responsibly. Attackers face none of that friction.

Casey notes that the answer is not fewer guardrails. Every capability useful for defense is equally available for offense, and pretending symmetry exists doesn’t help anyone. What does help, in his view, is closing the adoption gap on the defensive side while imposing real cost on offensive misuse. Lab-level access governance, he concludes, is the actual fight. And one that belongs squarely in a room like this.

Still unsolved: what stands in for patching

Is there a real alternative to patching being tried right now for CI operators caught in this gap? Honestly, he says, this is still unsolved – and he’d rather say that than pretend otherwise.

Segmentation, kill switches, and virtual patching are all real and all do real work, but he treats them as compensating controls rather than solutions. They’re reached for when patching isn’t possible, not instead of it, leaving the underlying gap open beneath them. Segmentation earns its keep by slowing propagation and containing blast radius rather than preventing intrusion. Kill switches are manual and depend on a human noticing first. Virtual patching buys time, but leaves noisier logs and a more fragile ruleset behind.

The binding constraint, in his view, isn’t cost or ignorance – it’s that downtime is the primary risk for these operators, which makes touching a running system to fix a confidentiality or integrity problem exactly the trade they’re built to refuse. The industry has been telling critical infrastructure to solve this for fifteen years without success, he argues, because the task kept being framed as accepting the industry’s risk trade rather than the operator’s own.

When an asset can’t be patched, who has the authority to decide what happens next? Who decides the timeline for accepting that risk? Who is financially or legally on the hook if that decision goes wrong? Right now, nobody clearly owns that decision – which is itself the actual failure, more than any missing tool.

What he’d actually do is stop optimising for prevention that can’t realistically be achieved, and assume compromise where instrumentation isn’t possible. The open question, in his framing, isn’t which control wins – it’s who owns the risk decision when an asset can’t be patched, on what timeline, and who carries the cost. He goes further still, suggesting that non-cooperative defense may become not just viable but necessary for CI providers operating below the security poverty line.

To clarify, this is a mindset shift. Right now, the industry’s default posture toward critical infrastructure is: try to prevent the breach. But Casey just explained that for many CI operators, patching (the actual fix) is often impossible in practice, because touching a live system risks downtime, and downtime is the one thing these operators are built to avoid at all costs.

So his point is: if prevention genuinely isn’t achievable for a given asset, stop pretending it is and stop pouring effort into a goal you can’t hit. Instead, for any system you can’t actively monitor or instrument (i.e. you have no visibility into what’s happening on it), the better assumption is: treat it as already compromised. Not ‘it might get breached’ – assume it is breached, and plan your defenses (segmentation, containment, response) around that assumption rather than around the fantasy that you’ll catch and stop the intrusion at the door.

Responsibility without a targeter

If an AI system can scan an entire class of critical infrastructure equipment without a human directing it at any specific target, how should responsibility be assigned if something goes wrong downstream? He’s not a lawyer, he notes, but in practice the traditional answer combines ‘whoever had the intent that triggered the harm’ with ‘whoever is legally weak enough to be successfully sued or prosecuted”. Neither half, he argues, survives contact with an agent scanning an entire equipment class.

His own test to check where responsibility sits is to see who set the objective. Until an AI picks its own target, it remains a tool, and the operator stays accountable for it. He points to Anthropic’s documented case of a state-sponsored campaign in which a model ran 80 to 90% of tactical operations across roughly thirty targets – and still calls it a tool, because humans chose the victims.

He resists framing this as an exotic new AI problem. A human always did something: chose the model, removed the constraints, defined the goal, set the scope, or failed to bound it. The useful work, in his view, is designing the ‘break-glass’ rules before the fire – who holds the authority, past what threshold, with what scope, what logging, what liability, and what accountability when it goes wrong.

Does the norm against attacking infrastructure still hold?

When an attack has no clear human decision-maker behind it, does the expectation that states won’t damage each other’s critical infrastructure still hold up (according to the UN cyber norms)? The expectation holds, he says. Its enforcement mechanism doesn’t – and those are two different problems that keep getting answered as one.

What existed, he argues, was never a treaty. It was a set of unspoken agreements maintained through strategic ambiguity and rational self-interest, and for fifteen years they held. Attribution was the mechanism that kept them in place: you can’t maintain mutually assured destruction if you can’t tell who launched the missile.

AI degrades attribution badly, he continues, and campaigns run in any language, exploitation chains automated without distinctive tradecraft signatures. Put a genuinely autonomous actor at the end of that chain, and the question of ’ho decided’ may have no clean answer at all.

He reads this as an argument for repairing the norm, rather than retiring it. A norm nobody can currently enforce still shapes what states are willing to be seen doing, and that residual effect is worth defending. The work ahead, in his framing, is rebuilding attribution and accountability until the norm means something again – which is precisely what a process like this one exists to support.

Are the CRA’s reporting clocks the right fix?

Asked whether the EU Cyber Resilience Act’s 24-hour and 72-hour reporting timelines are adequate, or already out of date, he was careful to separate support for the regime from a critique of one of its mechanics. He’s been positive on the CRA – its passage, alongside DORA, is what has most accelerated the policy-driven spread of secure-by-design as an accountability principle – and he doesn’t want a timeline critique read as an argument against the regime itself.

The clocks themselves were never going to match the tempo of the attacker, and he doesn’t think they can. What they do is end the silence, and for that purpose he considers them sound.

His concern is the destination, not the speed. Article 11 has operators handing governments information about actively exploited, unpatched vulnerabilities within 24 hours, while the regulation says remarkably little about what happens to that information afterward. disclose.io joined a letter signed by more than 50 experts opposing that requirement in 2023, and pushed again in 2025 for flexibility on the 72-hour mitigation window. His own position on the underlying equities question is ‘disclose by default, retain by exception’ – ground the CRA doesn’t clearly occupy in either direction.

Capacity compounds the problem, he adds. NIST announced in April that it could no longer enrich most CVEs, following a 263% surge in submissions, with roughly 29,000 backlogged. A reporting deadline is only ever as good as the capacity sitting behind it. September brings the first global mandatory exploited-vulnerability reporting regime – and his closing point is that the destination deserves as much attention as the clock.

Conclusion: the shift almost nobody is making

Asked what single shift in mindset this moment demands that almost nobody is actually making, he pointed to the OODA loop (Observe, Orient, Decide, and Act). It has collapsed, he argues, and almost nobody has changed what they own because of it. Observe, orient, decide, act has been compressed by AI to a width humans literally can’t fit inside. The work ahead is rethinking which steps a human retains, and accepting that others get delegated to systems still being learned to trust – work he stresses is not a tooling exercise, and one very few organisations have actually done.

Underneath that sits a bigger claim. The industry keeps treating this as a technical problem, he argues, when crime and the opportunities for crime are deeply rooted in sociology – cybercrime is a sociological problem amplified by technology, with technology merely providing the theatre. Every question raised in this conversation, in his view, is really a question about incentives and accountability, and the field keeps answering them with products instead.

He calls this moment the ‘slopdemic’ – the last cheap dress rehearsal available before the incentives themselves need fixing. As he puts it: he’d spend it on the incentives.