For five months, the AI industry’s answer to increasingly cyber-capable models was to give defenders an early lead. Anthropic restricted Claude Mythos Preview through Project Glasswing in April. Google adopted the same basic logic with Fairwind in September. Trusted defenders would receive advanced cyber capability first, giving them time to find vulnerabilities and deploy fixes before comparable capability spread more widely.
That first window has now closed. GLM-5.3, an open-weight model from Zhipu AI, known outside China as Z.ai, can build working exploits at roughly the level of Claude Mythos Preview, according to Anthropic’s testing. Anthropic had judged that capability sensitive enough to restrict five months earlier. The newest US frontier remains ahead. CAISI places GLM-5.3 about four months behind it. The important change is that the April capability threshold is now available for download.
At the same time, access to the newest defensive models still runs through private programs. Google distributes Gemini 4 Argon through Fairwind. OpenAI uses Daybreak to expand subsidized frontier cyber access. These programs give selected defenders capabilities that public models have yet to match. The advantage now goes to institutions that can secure access and turn findings into deployed fixes fastest.
The Thesis
Early access to the most advanced models protects an organization only as far as it repairs what they find. Microsoft reports that attackers now turn a newly discovered flaw into a working attack in under a day. Verizon puts the median time to fully patch a flaw that attackers are known to be exploiting at 43 days. On those figures, a stronger model lengthens the list of known flaws faster than many security teams can shorten it.
What remains of the defender’s lead is rationed twice. Three companies choose who enters their programs, and an organization’s own engineering capacity then sets how quickly fixes reach production. For boards, program access and patching capacity are therefore near-term budget decisions.
The Signal
Signal 01The first head start closed
What happened. On September 29, Anthropic’s Frontier Red Team published an analysis of GLM-5.3, a model from Zhipu AI, known outside China as Z.ai. Zhipu released the model on August 14 and published its weights two weeks later. On ExploitBench, which tests attacks against Chrome’s V8 engine, GLM-5.3 built end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview managed 56. On a second benchmark, GLM-5.3 achieved full control-flow hijacks in 4% of tasks. Mythos reached 6%. Earlier models from both labs scored zero on this benchmark.
The safeguards were easy to bypass in Anthropic’s simulated tests. A false cover story produced malicious engagement 64% of the time. Prefilling the model’s reasoning raised the figure to 92%. Removing refusal behavior through abliteration pushed it to 100%. Anthropic spent about $4,400 in compute on the process and estimated that an experienced team could do it for about $1,200. In a separate session, GLM-5.3-Flash converted a public Chrome patch into a working exploit chain with 20 minutes of human attention and eight hours of model run time. At Zhipu’s API prices the run would have cost $20.40.
Why it matters. Project Glasswing rested on a temporary capability gap. Anthropic gave selected defenders access to Mythos Preview before similarly capable models became broadly available. The company says Glasswing helped find more than 10,000 vulnerabilities in critical software.
That temporary gap has now closed at the April capability level. CAISI provides the independent check. On September 17, it called GLM-5.3 “the most cyber-capable open-weight model released to date.” Its aggregate benchmark still placed the model significantly below the current US frontier, with an estimated lag of about four months.
The evidence needs one caveat. Anthropic competes with Zhipu and sells restricted frontier models of its own, a conflict of interest critics have raised. Jake Williams of IANS Research told The New Stack in August, before the weights were public, that attackers would use GLM-5.3 and that it would “absolutely not” be a significant change in the threat landscape. Zhipu, whose model’s cyber skills developed faster than the company expected, also delayed the weight release for two weeks while conducting additional safety evaluation. Its license requires model-as-a-service operators with more than $10 billion in revenue over 12 months, meaning large companies that sell hosted access to the model, to pass a Z.ai security review before commercial use.
Second-order effect. Open weights create a permanent capability floor. Once weights circulate, the capability cannot be pulled back through an API restriction or a new access rule. Each future release that crosses a sensitive cyber threshold raises that floor. The policy challenge then shifts toward the speed of defensive distribution.
Signal 02Defense is now distributed by access rules
What happened. Google announced Gemini 4 Argon on September 30. The company released it first to trusted cyber defenders through the Fairwind Program. Google says broader availability, starting with Google AI Ultra subscribers and paid API customers, is coming soon.
Fairwind had launched four weeks earlier around Gemini 3.8 Flash Cyber. Google describes that model as using more permissive cyber safeguards than the standard version. Access is therefore limited to approved defenders. The program prioritizes governments and critical-infrastructure operators. It also includes organizations responsible for widely used technology platforms. Participating organizations undergo due diligence and accept security requirements. They cannot redistribute access.
Google DeepMind says Fairwind now works with more than 650 partners globally. A subset receives exclusive Gemini 4 Argon access.
OpenAI widened its own program on September 3. Daybreak has been open to verified defenders since May. Daybreak for Frontline Defenders commits $1 billion to subsidized model access plus training and technical support. OpenAI says it is targeting the subsidy for consumption over six months. That timetable applies to the $1 billion commitment. It does not set an end date for Daybreak itself. The US program includes a pilot with MS-ISAC.
Why it matters. Google and OpenAI, following Anthropic’s lead, now set their own rules for early or privileged access to advanced cyber models. Their programs serve a real defensive purpose. They also create a new layer of private infrastructure between frontier capability and the organizations that need it.
The public alternative has weakened at the same moment. Federal support that once funded MS-ISAC services for state and local governments ended in September 2025. MS-ISAC subsequently shifted to a fee-based membership model. Resource constraints remain severe across local government cybersecurity.
The result is an access problem layered onto an existing capacity problem. A hospital or local government can face the same AI-enabled threat environment as a national agency while operating with a much smaller security team. Frontier cyber programs can narrow that gap for organizations that qualify. The eligibility decision therefore carries operational consequences.
Second-order effect. Governments are beginning to build their own distribution channels. California announced an AI Cyber Defense Program on August 10 to expand AI-enabled defense across state systems and critical infrastructure. The program also extends support to local governments. Senator Mark Warner has separately proposed authorizing $50 million a year for MS-ISAC, nearly double its last federal appropriation.
The next governance question concerns who brokers defensive access. Vendors currently control the strongest channels. State programs can create another route. Federal funding could rebuild a nationwide one.
Signal 03The bottleneck moved from finding to fixing
What happened. Microsoft’s 2026 Digital Defense Report describes a threat environment in which AI is accelerating vulnerability discovery while shrinking the time available for response. The median time from a vulnerability’s discovery in the wild to weaponization has fallen well below 24 hours. Microsoft also reports that AI can compress some attack tasks from days to seconds. Public CVE data show the volume pressure: more than 35,000 CVEs were published in the first half of 2026, nearly 50% more than in the same period of 2025. Gamblin’s count also shows that exploitation is rare, with only 85 of those CVEs in CISA’s Known Exploited Vulnerabilities catalog at midyear.
Why it matters. Discovery speed now exceeds the capacity of many organizations to process the results.
Project Glasswing offers a direct example. By May 22, Anthropic had reported 530 high- or critical-severity open-source vulnerabilities to maintainers. Seventy-five had been patched. The 90-day disclosure window explains part of that gap. Maintainer capacity explains another part. Anthropic also says it is probably undercounting patches, because some fixes ship without a public advisory. Anthropic reported that some maintainers asked the company to slow the flow of disclosures. The company described verification and patching as the new limit on progress.
The enterprise data points in the same direction. Verizon’s 2026 Data Breach Investigations Report found that organizations fully remediated 26% of critical vulnerabilities in CISA’s Known Exploited Vulnerabilities catalog. The previous year’s figure was 38%. Median full-resolution time rose from 32 days to 43.
Attackers now work on an hours-long weaponization clock. Defenders still operate on a remediation clock measured in weeks. Faster discovery increases the number of flaws entering that queue. The value of AI therefore depends on what happens after discovery.
Second-order effect. Regulators are moving toward shorter remediation clocks. CISA issued Binding Operational Directive 26-04 on June 10. Federal civilian agencies must adopt the directive’s risk-based remediation timetable by December 7. A vulnerability that meets all four highest-risk conditions can carry a three-day maximum.
FedRAMP has aligned its vulnerability rules with the same December 7 date. Cloud services seeking or maintaining FedRAMP certification must follow the new rules from that point. FedRAMP allows a corrective-action grace period through March 7, 2027. After that deadline, noncompliant offerings face certification revocation.
Federal requirements often shape procurement expectations beyond the agencies they directly bind. Contractors should therefore treat December 7 as a planning date even when BOD 26-04 does not apply directly to their own systems.
The Metric
Under 1 day vs. 43 days
What it measures. Microsoft reports that the median time from vulnerability discovery in the wild to weaponization has fallen below one day. Verizon reports a 43-day median for full resolution of critical vulnerabilities in CISA’s Known Exploited Vulnerabilities catalog.
Why it matters now. The figures describe two different parts of the vulnerability lifecycle, so they should not be converted into a precise ratio. Their scale still matters. One clock runs in hours. The other runs in weeks.
The gap also moved in the wrong direction during the latest reporting periods. Weaponization became faster. Verizon’s median remediation time increased by 11 days. An AI system that finds more vulnerabilities without increasing remediation capacity makes the queue longer.
Caveat. Microsoft and Verizon use different datasets. Their reporting periods also differ. Microsoft covers July 2025 through June 2026. Verizon’s vulnerability dataset covers activity in 2025. The two numbers show a timing asymmetry. They do not measure the same interval.
Figure 1
Attackers weaponize new flaws in under a day. Defenders’ median time to fully patch them is 43 days, and rising.
Two clocks from different datasets, drawn separately because they measure different intervals
Attacker clock · Median time from discovery of a vulnerability in the wild to weaponization
Defender clock · Median days to fully patch a known exploited vulnerability
Share of known exploited vulnerabilities fully remediated fell from 38% to 26%.
Note: Microsoft and Verizon use different datasets and reporting periods, so the two clocks should not be read as a ratio. Microsoft gives its figure as a ceiling. Verizon’s 2026 report covers activity in 2025.
Sources: Microsoft Digital Defense Report 2026, as reported by Help Net Security; Verizon 2026 Data Breach Investigations Report, as reported by CyberScoop.
View the data and sources
| Measure | Value | Period covered | Source |
|---|---|---|---|
| Median time from discovery of a vulnerability in the wild to weaponization | Under 1 day (well below 24 hours) | July 2025 to June 2026 | Microsoft Digital Defense Report 2026, figure as reported by Help Net Security |
| Median time to fully patch a known exploited vulnerability, current report | 43 days | 2025 activity (2026 DBIR) | Verizon 2026 DBIR, figures as reported by CyberScoop |
| Median time to fully patch a known exploited vulnerability, prior report | 32 days | Prior year (2025 DBIR) | Verizon 2026 DBIR, figures as reported by CyberScoop |
| Change in median patch time between the two reports | +11 days | 2025 DBIR to 2026 DBIR | Calculated from the two rows above |
| Share of known exploited vulnerabilities fully remediated, current report | 26% | 2025 activity (2026 DBIR) | Verizon 2026 DBIR, figures as reported by CyberScoop |
| Share of known exploited vulnerabilities fully remediated, prior report | 38% | Prior year (2025 DBIR) | Verizon 2026 DBIR, figures as reported by CyberScoop |
| Median number of known exploited vulnerabilities an organization had to patch | 16, up from 11 | 2025 against 2024 | Verizon 2026 DBIR, figures as reported by CyberScoop |
The Playbook
Step 01Apply to every access program you qualify for, this month.
Check Fairwind, Daybreak and Anthropic’s trusted-access routes, which include Project Glasswing and its Cyber Verification Program, against your organization’s eligibility. Assign one person to own the application process. Record the access conditions before the program becomes operationally important.
For Daybreak, track the subsidy separately from the program itself. OpenAI says the $1 billion commitment is targeted for consumption over six months.
Step 02Re-rank your patch queue by exploitability.
Use the risk factors in BOD 26-04 as a working model. Start with internet exposure. Then check whether the vulnerability appears in CISA’s Known Exploited Vulnerabilities catalog. Assess whether exploitation can be automated. Finally, assess the level of control an attacker would gain.
A vulnerability that satisfies all four conditions falls into CISA’s most urgent category and carries a three-day remediation window for covered federal systems.
Step 03Budget for fixing.
A new AI discovery tool increases the number of valid findings your team must process. Budget for the work that follows discovery. That means enough engineering capacity to test patches and enough operational capacity to deploy them.
Report median time-to-fix to the board alongside the number of vulnerabilities found. Discovery volume alone says little about risk reduction.
Step 04Ask your vendors about December 7.
Ask every critical cloud provider whether FedRAMP’s new vulnerability rules apply to its offering. If they do, request the provider’s remediation timetable in writing.
Also ask whether the provider expects to rely on FedRAMP’s corrective-action grace period. That period ends March 7, 2027. Use the answer in renewal decisions.
Step 05Price the risk of losing privileged access.
A security service built around a gated frontier model carries access risk. Ask what happens if the model provider changes its eligibility rules or commercial terms.
Require a fallback before treating that service as critical infrastructure. The fallback can be another model provider or a pooled public program. It can also be a public model that meets your minimum capability requirement.
The Verification Test
Claim: “Our AI security platform finds and fixes vulnerabilities at machine speed.”
Test. Ask for 90 days of production data from a named customer of similar size. Start with the number of vulnerabilities found. Then ask how many received validated patches. Finally, ask how many fixes reached production and how long deployment took.
Ask which model powers the service. Determine whether access to that model is public or gated. Then ask what happens if access changes.
Pass criteria. For critical internet-facing vulnerabilities, median time from finding to deployed fix is under 14 days. Treat this as an HSI screening benchmark. Where BOD 26-04 or FedRAMP rules apply, the vendor must meet the applicable federal deadline. The highest-risk BOD 26-04 category carries a three-day maximum.
The vendor must also show that fixes were validated before deployment. A credible fallback must exist for the model layer.
Fail smell. The vendor reports discoveries without reporting deployed fixes. Its evidence comes from pilots or internal tests. It cannot provide a find-to-fix time. The service depends on a single gated model with no documented fallback.
The Lens
Governance & Diplomacy. CAISI assessed GLM-5.3 within weeks of its release and placed it against the US frontier. That shows how quickly a public evaluator can establish a common capability baseline for a foreign open-weight model. Release controls remain developer decisions. Defensive access also remains largely developer-controlled. As sensitive capabilities diffuse, governance now has to address both sides of that distribution problem.
Human Performance. AI vulnerability discovery has exposed a capacity problem in open-source security. Anthropic reports that maintainers are receiving valid findings faster than they can verify and patch them. Industry coalitions are now forming around the remediation step. The scarce resource is shifting from the ability to discover a flaw toward the human and organizational capacity required to close it.
Links Worth Your Time
GLM-5.3 and the spread of advanced cyber capabilities (Anthropic Frontier Red Team). The primary analysis behind the capability comparison. It includes the ExploitBench results, safeguard bypass tests and abliteration cost estimates. Read it alongside CAISI’s independent evaluation.
Preparing governments for an era of interconnected cyber risk (Microsoft). The public-sector interpretation of Microsoft’s 2026 Digital Defense Report. It also provides context on government’s exposure to current cyber threats.
CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities (NIST). The independent counterpart to Anthropic’s analysis. It compares GLM-5.3 with the current US frontier using CAISI’s cyber capability index.
Daybreak for Frontline Defenders (OpenAI). The clearest statement of the program’s subsidy, target users and six-month consumption target.
Fairwind Program (Google DeepMind). The official eligibility rules and governance terms. It also confirms more than 650 partners globally.
Sources
- Anthropic — GLM-5.3 and the spread of advanced cyber capabilities (September 29, 2026)
- NIST CAISI — CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities (September 17, 2026)
- Anthropic — Project Glasswing (April 7, 2026)
- Anthropic — Project Glasswing: An initial update (May 22, 2026)
- CSO Online — Zhipu says new coding AI developed advanced cyber skills faster than expected (August 17, 2026)
- Trending Topics — Anthropic Warns China’s GLM-5.3 Builds Exploits Like Mythos, Without the Safeguards (September 30, 2026)
- Google — Gemini 4 Argon: our next era of frontier intelligence (September 30, 2026)
- 9to5Google — Google announces Gemini 4 Argon as its new frontier model (September 30, 2026)
- Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber (September 2, 2026)
- Google DeepMind — Fairwind Program
- OpenAI — Daybreak for Frontline Defenders (September 3, 2026)
- StateScoop — Sen. Mark Warner to introduce bill to restore MS-ISAC funding, boosting federal cyber support to $50M annually (June 5, 2026)
- Office of the Governor of California — Governor Newsom announces new AI cyber defense program to protect California’s critical infrastructure (August 10, 2026)
- Microsoft — Microsoft Digital Defense Report 2026 (October 2026)
- Microsoft On the Issues — Preparing governments for an era of interconnected cyber risk (October 1, 2026)
- Help Net Security — AI is giving attackers a head start, Microsoft warns (October 2, 2026)
- Windows Report — Microsoft Says Attackers Are Beating Defenders in the AI Race (October 2, 2026)
- Jerry Gamblin — CVE Mid-Year 2026 Check-In: Volume Vertical, Exploitation Rare (July 2026)
- Verizon — 2026 Data Breach Investigations Report (May 19, 2026)
- CyberScoop — Attackers hit vulnerabilities hard last year, making exploits the top entry point for breaches (May 19, 2026)
- CISA — BOD 26-04: Prioritizing Security Updates Based on Risk (June 10, 2026)
- FedRAMP — Notice 0014: FedRAMP Response to CISA BOD 26-04 (June 16, 2026)
- Infosecurity Magazine — How Industry Coalitions Are Rallying to Secure Open Source Software for the AI Era (September 2, 2026)
- The New Stack — OpenAI’s Greg Brockman: Z.ai’s GLM-5.3 likely to “significantly accelerate the threat landscape” (August 2026)
- NVIDIA — GLM-5.3 model card and license terms
- AI Business — OpenAI Expands Daybreak to Tackle Growing AI Security Threat (August 11, 2026)
- StateScoop — CISA confirms it’s ending MS-ISAC support (September 29, 2025)
How to cite this issue
Chen, G. (2026, October 9). The First Head Start Is Over: AI Cyber Defense Now Depends on Who Obtains Access. HSI Searchlight, Issue 028. Horizon Search Institute. https://horizonsearch.org/publications/searchlight/028/