arXiv · Jul 17, 2026

Harmonizing AI Safety Thresholds

An arXiv paper proposes a method to harmonize safety thresholds across frontier AI companies in three risk areas: cyber misuse, biological misuse, and automated AI R&D, addressing inconsistencies that hinder verification and comparison.

What happened

Frontier AI companies publish widely varying capability thresholds, making it difficult for third parties to verify if thresholds are breached or compare requirements across companies. Lack of common minimum thresholds may lead to inconsistent risk mitigation and a race to the bottom in safety standards. This study develops a method to derive harmonized thresholds in three risk areas: for cyber and biological misuse, it uses expected harm as a key primitive with explicit risk modeling; for automated AI R&D, thresholds are based on observed AI progress rates rather than expected harm. The analysis extends prior work and identifies empirical gaps and limitations.

Technical significance

The method uses expected harm as the core metric for misuse risk thresholds, with explicit modeling of risk pathways and model release conditions; for automated AI R&D, it innovatively uses AI progress rates as the threshold basis instead of harm assessment.

Industry impact

Current lack of unified safety thresholds among frontier AI companies may lead to regulatory arbitrage and declining safety levels; the industry urgently needs comparable benchmarks to harmonize safety practices.

What to watch

Subsequent observation will show whether regulators or industry consortia adopt such harmonized thresholds, and whether companies adjust their published safety frameworks to align with common standards.

Decision value

Harmonized thresholds help reduce compliance costs, enhance public trust, and provide a standardized basis for commercial services such as AI safety audits and insurance.

Evidence