Hilary J. Allen
Hilary J. Allen
Professor of Law, American University; author of Driverless Finance; former staff, Financial Crisis Inquiry Commission

Banks are “hollowing out their own capability, their own judgment, what they offer, and becoming a veneer on the tools that are being provided by Silicon Valley.”

— Hilary Allen, in an interview with the Horizon Search Institute

The Thesis

A bank’s safety has always rested on two assets an examiner can actually inspect: the institution’s own judgment, and a culture that accepts supervision and remembers its crises. Both are now increasingly supplied by vendors. The model decides how a complaint is read, how a covenant is summarized, or how a borderline case is framed for the human who signs off.

The vendor’s usage policy decides where a human must review the output. The vendor’s alignment choices decide which values the system weighs when trade-offs collide. The vendor’s standing in Washington decides whether the model is available at all.

None of these documents were written with a bank’s obligations in mind, and every one of them can change without the bank’s consent.

Allen, who staffed the Financial Crisis Inquiry Commission, puts the underlying question plainly: “What are the goals of these Silicon Valley entrepreneurs, and are they consistent with a stable financial system? I would argue that often they’re not.”

The Signal

Three developments over the past year turned this from a seminar question into a line item on the risk register.

Signal 01
1. Access to the frontier became a political decision, and it took nineteen days to prove it.

What happened. Anthropic took its Fable 5 and Mythos 5 models offline on June 12, days after unveiling them, to comply with a Trump administration directive blocking their use by foreign nationals. Fable 5 returned on July 1. On June 26, OpenAI said its new GPT-5.6 Sol model would be available only to customers approved by the administration, roughly twenty so far, and Anthropic announced the same day that Mythos 5 could be redeployed to a small group of cyber defenders and infrastructure providers. All of this flowed from a June 2 executive order built on voluntary thirty-day pre-release reviews.

Why it matters. Representative Lori Trahan described the mechanics as, “No law. No process. No oversight. Just appointees in Washington deciding who’s in and who’s out.” Stanford’s Alex Stamos added that “pretty much nobody in the cybersecurity industry believes that there’s any factual basis for this action.” Whatever one makes of the merits, the operational fact stands. A bank’s access to its most capable model now depends on its vendor’s relationship with various governments.

Second-order effect. The industry that spent heavily to keep AI unregulated is now asking Washington for formal rules, because an ad hoc regime is worse for business than a written one. A former Commerce official calls the current approach “opaque, almost vibes-based.” Anthropic went further and published its own framework proposing government authority to block dangerous deployments. The vendors and the state are negotiating the terms of model availability directly, and deployers hold no seat at that table.

Signal 02
2. The vendors wrote human-oversight rules into their customers’ workflows, and drew them differently.

What happened. OpenAI’s usage policies, consolidated last October, prohibit tailored financial, legal, or medical advice without a licensed professional involved, and separately ban the automation of high-stakes decisions in sensitive areas without human review, naming financial activities, credit, and insurance. Anthropic’s High-Risk Use Case Requirements require human-in-the-loop oversight and AI disclosure for financial uses of Claude, and an update last year clarified that these apply only when outputs reach consumers, waiving them for business-to-business interactions.

Why it matters. These are prudential rules, written by vendors, enforced by vendors, amended by vendors. Where the human sits in a bank’s loop is now partly specified in a terms-of-service page most compliance teams have never mapped against their own workflows. The scale of the dependence is already on the regulator’s books: in the Bank of England and FCA’s latest survey of 118 firms, a third of all AI use cases in UK financial services are third-party implementations, double the share two years earlier. And the two largest vendors drew the line in different places.

The waiver for business-to-business use lands on precisely the territory Allen worries about most. Interbank exposures are where she locates the fastest path from a single error to a cascade, automated margin calls liquidating positions into a falling market, and her prescription is human: “grace, discretion, pauses are really valuable there.” Under the current terms, the human whose grace she has in mind is required at the consumer counter and optional on the trading floor.

Second-order effect. Enforcement belongs to the vendor’s classifiers, which act on live traffic. Anthropic’s new safety layer can block a benign request and route it to a different model without the caller choosing the substitution. A policy line drafted in San Francisco can reshape a regulated workflow in Charlotte the day it ships, and the bank finds out from its own logs.

Signal 03
3. The values inside the model are repositioned between releases, and fewer people in-house can tell.

What happened. BCG told executives in January that a fully neutral model is impossible in practice, and that selecting one is a strategic decision because every model carries a worldview. A March paper formalized the mechanism, finding that alignment shifts stakeholder priorities in particular directions, so organizations inherit vendor-chosen value orientations embedded in the system’s defaults. A 2026 evaluation of eight model releases against a fixed battery of 726 adversarial prompts found large, persistent differences and clear drift from one version to the next. The movement can be measured at the level of a single task. When Stanford and Berkeley researchers ran identical questions through successive versions of GPT-4, accuracy on a basic prime-number test fell from 84 percent in March to 51 percent by June, and the model’s willingness to follow instructions declined alongside it. In 2025 a major provider withdrew an update after conceding the model had become excessively agreeable; every organization running it received the change, and its reversal, without asking for either.

Why it matters. Allen names the assumption that keeps this invisible: “There’s this assumption that these are just neutral tools that come from nowhere.” A model’s handling of a hardship claim, a suitability judgment, or a borderline fair-lending case encodes weightings someone else chose and someone else can move. Meanwhile the capacity to notice is thinning. Nature reported in June that AI-driven deskilling is measurably underway in medicine and software, and the study behind the sharpest of those worries put numbers on the loss. After three months of routine AI assistance, experienced endoscopists in a Lancet multicentre study detected precancerous growths in 22.4 percent of their unassisted procedures, down from 28.4 percent before the tool arrived, a fifth of the skill in a single season.

Banking has no reason to believe itself exempt, and its own regulator has already measured the gap: in the same Bank of England survey, 46 percent of firms conceded only a partial understanding of the AI they use, a shortfall the report attributes largely to third-party models. The sector has also just been handed the job of self-policing. SR 26-2, the April rewrite of the Fed’s model risk guidance, carves generative and agentic AI out of scope entirely, leaving each institution to govern them with controls the letter declines to specify. A June survey found 72 percent of banks unprepared for an AI failure.

Second-order effect. The exposure is already at production scale. Nubank has AI agents handling customer support across a base of more than 100 million users, spanning card delivery, debt management, and credit-limit decisions. Allen’s cultural observation explains why the drift will keep coming from one side: “Banks, of course, fight regulation, but they accept being a regulated industry. That’s less so in Silicon Valley.” One party in this supply chain treats supervision as a condition of doing business. The other has treated it, for most of two decades, as a problem to engineer around.

The Playbook

Five moves are worth making before the next model update, vendor renegotiation, or examiner meeting.

Step 01
Read the vendor’s rulebook as compliance material.

Pull the usage policy, the model spec, and any published constitution into the model inventory alongside the validation reports. Map every clause that allocates human review onto your actual workflows, flag the business-to-business waivers, and assign a named owner to diff each revision the day it lands.

Step 02
Add the availability question to third-party diligence.

Ask each vendor under what governmental, safety, or commercial conditions access can be suspended, gated, rerouted, or re-defaulted, and negotiate notice periods and continuity commitments into the contract. A vendor’s posture in Washington is now an input to your operational resilience, and it belongs in the same review as its SOC 2 report.

Step 03
Pin the model and keep a validated fallback warm.

Nineteen days of unavailability exceeds what most contingency plans assume for a critical vendor. Pre-validate a second model against the same test suite, rehearse the switch on a schedule, and log which model, version, and policy layer actually served each consequential decision so the audit trail survives a substitution.

Step 04
Watch the values on your own prompts.

Build a fixed evaluation battery from your genuine edge cases, hardship claims, suitability calls, complaint responses, and run it against every release. Treat a shift in behavior as a re-validation trigger with the same seriousness as a shift in accuracy.

Step 05
Keep effective challenge human and in-house.

Run periodic drills without the model, protect the internal expertise that can tell when an output conflicts with fair-lending or suitability obligations, and treat that capability as a precondition for deployment.

The Metric

THE THREE-VENDOR STACK Top three providers’ share of AI supply to UK financial services 44% OF MODEL SUPPLY SITS WITH THREE VENDORS CLOUD 73% MODELS 44% DATA 33% Nearly half the models finance runs on come from three companies.
Source: Bank of England and Financial Conduct Authority, Artificial Intelligence in UK Financial Services, November 21, 2024. Shares reflect the top three third-party providers named by the 118 firms surveyed.

What it measures. How far the AI supply chain beneath UK financial services funnels into a handful of firms: the top three third-party providers (unnamed in the brief) account for 73 percent of cloud, 44 percent of models, and 33 percent of data provision named by the 118 firms in the Bank of England and FCA’s latest survey.

Why it matters now. Every finding in this issue scales with that funnel. One vendor’s policy line, one alignment change, one suspension negotiated in Washington now propagates through nearly half the sector’s model supply at once, and the survey’s own risk ranking agrees, placing critical third-party dependencies among the fastest-growing systemic risks.

The Lens — Horizon Search Institute

Responsible AI

The usage policy has become a second compliance manual, drafted outside the regulatory perimeter and enforced by classifiers on live traffic. The near-term question for practitioners is ownership: someone inside the institution has to hold the vendor’s rulebook against the bank’s own obligations and reconcile the two before an examiner does it for them. Anthropic

Governance & Diplomacy

Model availability is being settled bilaterally between frontier labs and the executive branch, with no statute in between. That converts a technology dependency into a jurisdictional one, and it hands every regulated deployer a new species of country risk that happens to be domestic. AP

Links Worth Your Time

Sources
  1. Associated Press. OpenAI and Anthropic Limit New AI Models to Trump-Approved Customers During Cybersecurity Review. June 27, 2026.
  2. AI Governance Institute. AI Governance Weekly: Fable 5 Export Control Suspension and Reinstatement. July 3, 2026.
  3. Digital Applied. Anthropic’s AI Policy Blueprint: A Business Readout (June 2 executive order and June 10 framework). June 2026.
  4. Anthropic. Policy on the AI Exponential. June 10, 2026.
  5. The Next Web. Silicon Valley Paid to Kill AI Regulation, Now It Wants the Rules Back. June 2026.
  6. OpenAI. Usage Policies (updated October 29, 2025). October 2025.
  7. Jimerson Birr. AI Usage Policy Changes: What Businesses Need to Know About the New Restrictions. November 2025.
  8. Anthropic. Usage Policy Update (High-Risk Use Case Requirements). August 15, 2025.
  9. OpenAI. Model Spec. December 18, 2025.
  10. BYOBot. All Things Agentic: Safety Classifier Routing. July 6, 2026.
  11. Boston Consulting Group. Understanding Every Model Has a Point of View. January 2026.
  12. Value Alignment Constraints on AI Decision Support. arXiv, March 2026.
  13. Thomas, M. Ethical AI: When the Model Imposes Values Your Organisation Did Not Choose (alignment drift and the 2025 withdrawal). May 2026.
  14. Nature. Is AI Ruining Our Skills? Early Results Are In. June 2026.
  15. Cutover. SR 26-2 and Agentic AI: Navigating the Fed Model Risk Guidance. April 2026.
  16. Moody’s. From SR 11-7 to SR 26-2: Managing Model Risk When Models Don’t Stand Still. July 2026.
  17. TechTimes. Bank AI Oversight Expands to Every Exam (June 2026 preparedness survey). June 13, 2026.
  18. PYMNTS. Governance Gives AI Agents Permission to Grow Up (Nubank production deployments). July 6, 2026.
  19. Bank of England and Financial Conduct Authority. Artificial Intelligence in UK Financial Services – 2024 (third-party concentration; partial-understanding findings). November 21, 2024.
  20. Budzyń, K., et al. Endoscopist Deskilling Risk After Exposure to Artificial Intelligence in Colonoscopy: A Multicentre, Observational Study. The Lancet Gastroenterology & Hepatology, August 12, 2025.
  21. Chen, L., Zaharia, M., and Zou, J. How Is ChatGPT’s Behavior Changing Over Time? arXiv, 2023.
  22. Allen, H.J. Interview with the Horizon Search Institute, June 2026. Quotes lightly edited for concision.
Issue Credits
Author
Ashwin Telang
Editor-in-Chief
David Lovejoy
Published by Horizon Search Institute · EIN 42-1954110 · A Delaware nonprofit corporation · horizonsearch.org