
Elon Musk is proposing a different kind of AI safety oversight. Instead of waiting for a global regulator, he wants the leading AI labs to test each other’s models.
Musk said top US AI companies and three or four leading Chinese AI companies should allow rivals to run a test harness on their models to evaluate safety, according to CNBC’s Techmeme-tracked report and related coverage from The Information. The idea is that serious risks would be exposed by adversarial peer review before a powerful system reaches the public.
The proposal fits Musk’s long-running view that AI safety needs more than company self-certification, but it also avoids the kind of centralised global body that some industry figures distrust. Rival labs would test one another because they have the technical skill and the incentive to find weaknesses.
The attraction of the idea is obvious. OpenAI, Anthropic, Google DeepMind, Meta, xAI and leading Chinese labs understand frontier models better than most governments do. If they can test one another’s systems under controlled rules, they may catch dangerous behaviours earlier than a distant regulator with limited technical access.
But there are obvious complications. Companies may not want rivals near their model weights, training details, prompts, internal tools or safety methods. Chinese and American labs may distrust each other. Open-source advocates may worry that the largest companies will become a private safety club that excludes smaller developers. Governments may also ask why companies should be allowed to decide among themselves what counts as safe.
Still, Musk’s idea lands at a moment when the AI industry is already discussing embedded evaluators, standards bodies and model audits. TechBooky has covered OpenAI’s support for outside evaluators for frontier AI models and the wider argument that AI rivals may need a speed limit. Musk’s version is more competitive. Let rivals try to break the model before everyone else has to live with it.
The hard part will be enforcement. A peer-review system only matters if companies are required to submit their strongest models, disclose serious findings and delay releases when problems are found. Without that, it risks becoming another safety ritual that sounds impressive but changes little.







