A new analysis suggests that weak AI safety regulation can backfire, producing AI products that are potentially more dangerous than those built in a completely unregulated environment. The paper, published Monday in the Proceedings of the National Academy of Sciences, uses game theory and theoretical economics to map the safety incentives across the AI supply chain. Its central claim is that regulation aimed only at downstream companies that apply AI in real-world settings gives general-purpose AI developers a strong reason to shirk their own safety duties. The result can be an end product that is less safe than if no rules had been written at all.
Regulating the wrong link
Researchers from Cornell and Carnegie Mellon built a formal model with two types of players. The first type are the general-purpose AI producers, such as OpenAI, Google, and Anthropic, who build large foundation models designed to be used in many different contexts. The second type are downstream specialists, the companies that take those models and embed them into products like AI-powered medical diagnostic systems or e-commerce customer service chatbots.
The intuitive approach, the authors argue, is to regulate the downstream specialists because their products have concrete and identifiable risks. A medical AI, for instance, visibly affects patients. A certain type of chatbot can be tested against customer service outcomes. But the paper warns this intuition is misleading.
When the government focuses on downstream companies and lets the developers of general-purpose systems off the hook, the developers tend to cut corners on safety measures such as third-party audits and red-teaming. They know that a downstream specialist will still have to satisfy regulators before a product reaches the market. So they pass the safety burden down the line. The study’s lead author, Benjamin Laufer, describes this as a “free-riding” behavior. He says the regulation acts as a tool for the general provider to offload the safety burden onto the downstream specialist.
Because the downstream specialist rarely has complete visibility into the model’s training data, architecture, or internal guardrails, its ability to catch problems is inherently limited. If the general-purpose developer has already depressed its own safety investments, the specialist is trying to patch risks it may not fully understand. The end product reaches the public with less cumulative safety verification than in a no-regulatory scenario, where the general provider might feel more responsible for voluntarily ensuring quality across the chain.
A game theory dilemma
To explain this dynamic, the researchers frame the relationship as a classic prisoner’s dilemma. In game theory, two rational decision-makers can cooperate or betray each other. If both cooperate, both receive the best possible outcome. If one cooperates and the other betrays, the cooperator ends up with the worst outcome. If both betray, both get a mediocre outcome. Since neither knows what the other will do, each often chooses to betray, producing a worse collective result than cooperation would have delivered.
In AI safety, cooperation means investing in audits, documentation, model evaluation, stress-testing, and ongoing monitoring. Betrayal means skimping on those costs while hoping the business partner absorbs the risk. Unless there is a way to guarantee that the other side will also invest, each side has an incentive to shirk.
The researchers find that strict, well-placed regulation changes that incentive structure. When both general-purpose producers and downstream specialists are required to meet meaningful safety standards, each side can trust that the other will also invest. This creates a cooperation equilibrium that improves safety and economic outcomes for everybody. In the model, the “sweet spot” exists when the regulator expects both sides to invest enough to meet a clear safety bar. The utility for each firm—defined as revenue share minus investment cost—is higher under this coordinated equilibrium than under a free-riding equilibrium.
Two camps in the AI policy war
The study enters a highly charged policy debate. On one side are anti-regulation technologists and people aligned with the Trump administration’s approach to AI governance. This camp sees stricter federal rules as unnecessary and believes the industry should be free to innovate as fast as possible. Their argument is often framed as a national security necessity: if American companies slow down to comply with a thick rulebook, China could take the lead in the global AI race. Some in this group also dismiss strict-safety supporters as doomers, or accuse them of attempting regulatory capture to advantage incumbent firms.
On the other side are unions, consumer advocates, journalists, and many AI researchers. They argue that market incentives are not enough to control a technology with enormous and potentially systemic consequences. They point to harms that range from AI-generated disinformation and algorithmic bias to environmental impacts caused by data centers and a feared wave of unemployment as AI tools displace workers. They often frame regulation as an essential margin of safety for society, even if some measures add costs to companies.
The new paper argues this framing is too binary. Safety and revenue do not have to be an either-or situation. Strong and well-placed regulation can mutually benefit all players by improving both the safety of the end product and the utility that general-purpose AI creators and downstream specialists get from their investment. The authors are not saying that every aspect of AI should be heavily regulated; they are saying that when regulation is written, it should be structured to align incentives along the entire supply chain.
Key findings from the study
- Weak regulation that targets only one group in the AI supply chain can backfire and reduce overall safety.
- General-purpose AI developers may deliberately invest less in safety if they know downstream companies will face regulatory accountability.
- The dynamic is a textbook prisoner’s dilemma; cooperation on safety is fragile without a rule that commits both sides to invest.
- Strict regulation that applies to both general model developers and downstream specialists can produce a better outcome for all.
- In the model, firm utility is defined as revenue share minus investment cost, and it can increase under a coordinated regulatory equilibrium.
What the model means for policymakers
Laufer stresses that a complex technology like AI involves a very complicated set of stakeholders, each with their own contribution to the technology. “People think of AI as a single object,” he said, “but actually AI involves a very complicated set of stakeholders and actors that each have their own contributions to the technology. To regulate in a thoughtful way, we need to consider the whole supply chain, not just a single provider or entity.”
The lessons from other industries reinforce that lesson. In aviation, safety regulations apply to aircraft manufacturers, parts suppliers, and airlines, because a failure at any stage can take down a plane. In the pharmaceutical market, regulators oversee drug developers, manufacturers, and prescribing practices, because the safety of a medicine depends on each step. In banking after the 2008 financial crisis, regulators discovered that mortgage originators and securitizers both had to be constrained; if only one side was supervised, the other could churn out hidden risk.
Artificial intelligence has a similar structure. A foundation model is a component that can be used in thousands of products. The same model can power a legal research tool, a hiring screening system, a criminal risk assessment, and a mental-health chatbot. If the base model is weakly aligned, the downstream product is fragile no matter how much app-level testing the deployer does. Conversely, a safe base model can still be deployed in an unsafe context if the downstream company does not worry about application-specific consequences. A supply-chain-wide regime addresses both sides of the problem.
The difficulty of defining safety
One difficulty the authors acknowledge is that “meaningful safety standards” for general-purpose AI are not easy to define. There is no universally accepted benchmark for a safe foundation model. Different applications have different thresholds, and capability risks change as models improve. Marketers and developers may disagree on what evidence is enough. Even a strict regulator would need a practical way to certify model auditing and take enforcement action when a model developer falls short.
The study does not attempt to solve that technical problem. Its contribution is more structural: whatever the standards are, they should be applied to both model developers and deployers so that neither side can shift the entire burden to the other. The paper also does not claim that all AI uses are equally risky. A low-stakes e-commerce chatbot probably deserves a different set of requirements than a hospital triage system. But the researchers’ formal rule remains the same: if a product is regulated, the underlying provider of the model should not be able to escape accountability simply because the final use case is being policed.
The study also pushes back on the idea that stricter rules must inevitably lead to slower innovation. If firms can coordinate on safety investments because a regulator enforces a baseline, they can avoid costly catch-up failures, recalls, and public distrust. A model that has been rigorously tested and well-documented may be more commercially valuable, not less. The authors emphasize that regulation is not always a tax on progress; it can be a coordination device that helps the entire industry produce better products.
For now, the debate in Washington and Silicon Valley remains polarized. But this research offers a middle path grounded in formal theory. Rather than asking whether AI should be regulated at all, the question should be how to design rules that reach every link in the supply chain, from the groups building the models to the firms putting them in front of users. Anything less may create the false appearance of safety while quietly making the world less safe than if regulators had stayed out of the way altogether.
Source: Gizmodo News