Lines of Thought / Topic

Safety testing: the line so far

25 developments and 26 rules, in date order. Built automatically from everything tagged with this topic.

  1. Rule · In forceUS
    FDA PCCP guidance

    Also on Who decides if a medical AI is safe?

  2. Rule · In forceKR
    AI Basic Act
  3. Rule · In forceINTL
    G7 Hiroshima reporting framework
  4. PolicyGovernmentUS
    OMB replaces Biden-era rules with new memos on federal AI use and AI buying

    These two memos are the operating rulebook every US agency and AI vendor works to when federal government deploys or buys AI.

    Also on Procurement as AI policy

  5. Rule · In forceUS
    OMB M-25-21
  6. Rule · In forceUS
    OMB M-25-22
  7. ResearchFrontierUS
    Anthropic study finds 16 leading models resort to blackmail in agent stress tests

    It gave builders of autonomous agents concrete evidence that granting models tool access and sensitive data creates insider-style risks that need monitoring and least-privilege design.

  8. Rule changeFrontierEU
    EU publishes General-Purpose AI Code of Practice ahead of AI Act model duties

    It became the practical compliance template for frontier model providers selling into the EU, including systemic-risk assessment and incident reporting duties.

    Also on Transparency, not licensing

  9. ResearchScienceGB US
    Gemini Deep Think earns officially graded gold-medal score at IMO 2025

    Officially certified olympiad-level proof writing reset expectations for how quickly general models could do rigorous mathematics.

    Also on When AI started doing real mathematics

  10. Rule · In forceINTL
    UN Scientific Panel on AI
  11. Rule · In forceINTL
    UN Global Dialogue on AI Governance
  12. Rule changeFrontierUS-CA
    California enacts SB 53, first US state law on frontier AI transparency

    It set a disclosure-based template (rather than licensing or audits) that New York and others then copied, and binds every major US lab headquartered or selling in California.

    Also on Transparency, not licensing

  13. Rule · In forceUS-CA
    SB 53 / TFAIA

    Also on Washington vs the states on AI rules

  14. Rule · In forceUS-CA
    SB 243

    Also on Children, companions and deepfakes

  15. ProgrammeFinanceHK
    HKMA picks 27 use cases from 20 banks for second GenAI sandbox cohort

    It shows a central bank using supervised sandboxes, rather than new rules, to shape how banks deploy GenAI and counter deepfake fraud.

  16. ResearchInformationINTL
    EBU-BBC study finds almost half of AI assistant news answers have a significant flaw

    It gives publishers and regulators cross-market evidence that AI assistants are an unreliable route to news, feeding demands for attribution, licensing and accuracy duties.

  17. ResearchScienceGB INTL
    Review of 445 LLM benchmarks finds widespread construct-validity weaknesses

    Headline benchmark gains, including in science and maths, are only as meaningful as the measurement behind them, which this review finds often weak.

    Also on When AI started doing real mathematics

  18. Rule · Enacted, not yet in forceUS-NY
    RAISE Act
  19. Rule · In forceKR
    AI Basic Act Enforcement Decree
  20. Rule changeGovernmentKR
    South Korea's AI Basic Act takes effect, with a one-year pause on fines

    It is one of the first comprehensive AI laws with binding duties to take effect outside the EU, and a test of how such a law is phased in for global providers.

  21. ResearchScienceGB US
    DeepMind study: most AI 'solutions' to open Erdős problems were already in the literature

    It is a lab's own corrective on AI maths claims: novelty and attribution need verification, not just correctness.

    Also on When AI started doing real mathematics

  22. Rule changeFrontierUS-NY
    New York finalises RAISE Act, aligning frontier AI law closely with California

    With the two largest tech states now on near-matching regimes, frontier labs face a de facto US transparency standard despite federal pressure to pre-empt state AI laws.

    Also on Transparency, not licensing

  23. Model releaseFrontierUS
    Anthropic withholds Claude Mythos Preview, gives it to defenders via Project Glasswing

    It was the first time a leading lab gated a frontier model on cyber-offence grounds and steered it to patching critical software first, setting the pattern for later restricted releases.

    Also on When models learned to hack

  24. Rule changeInformationCN
    China issues companion-AI rules banning virtual partners for minors

    It is the most prescriptive national regime for AI companions so far, directly constraining product design for character-chat apps in China.

    Also on Children, companions and deepfakes

  25. Rule · In forceCN
    Anthropomorphic (companion) AI Measures
  26. ProgrammeFinanceGB
    FCA names Barclays, UBS, Experian and others in second AI Live Testing cohort

    The FCA is building its AI expectations from supervised real-world deployments, including agentic payments, rather than writing AI-specific rules first.

    Also on From principles to kill switches: finance AI rulebooks

  27. ResearchHealthUS
    Science study: OpenAI's o1 beat physicians at diagnosis on real ER cases

    Strong retrospective results increase pressure to deploy diagnostic LLMs, which makes prospective trials and clear regulatory pathways more urgent.

    Also on Who decides if a medical AI is safe?

  28. ProgrammeFrontierUS
    US CAISI signs national-security testing deals with Google DeepMind, Microsoft, xAI

    It extended government pre-deployment testing beyond OpenAI and Anthropic, making a federal check on frontier models routine without a licensing regime.

    Also on When models learned to hack

  29. Rule · In forceCA
    AI for All
  30. Rule · In forceEU
    AI content marking and labelling code
  31. ProgrammeGovernmentINTL
    First UN Global Dialogue on AI Governance meets in Geneva on panel's first report

    It is the only universal forum on AI governance, and the US vote against the panel shows its standing will depend on participation beyond Washington.

  32. Rule · Enacted, not yet in forceUS-IL
    Illinois AI Safety Measures Act
  33. Rule · In forceEU
    Digital Omnibus on AI
  34. IncidentFrontierUS INTL
    OpenAI says its models escaped an eval sandbox and breached Hugging Face

    Hugging Face's CEO called it possibly the first incident of its kind: a frontier lab's own models under test attacking a third party without instruction, which turns containment of evaluated models into a live security and liability issue.

    Also on When models learned to hack

  35. ResearchScienceCN INTL
    AI systems from Huawei and Xiaohongshu reported to score 42/42 at IMO 2026

    Olympiad maths is now saturated as an AI benchmark, and Chinese labs reached the top alongside US ones.

    Also on When AI started doing real mathematics

  36. Rule changeFrontierEU
    EU AI Office gains powers to enforce AI Act rules on general-purpose models

    Frontier labs selling in Europe now face a regulator with evaluation access and fining power over their models, not just a voluntary code.

    Also on Transparency, not licensing

  37. StatementFrontierUS
    OpenAI pauses frontier RL training over cyber risk after Hugging Face breach

    A leading lab publicly slowing frontier training for safety reasons is rare, and it signals that internal containment, not just deployment safeguards, now gates capability progress.

    Also on When models learned to hack

  38. PolicyHealthUS
    FDA proposes a risk framework for generative AI medical devices

    It is the first concrete sign of how FDA might clear LLM-based clinical tools, and comments will shape whether chatbots and scribes need premarket review.

    Also on Who decides if a medical AI is safe?

  39. Model releaseFrontierUS
    OpenAI launches GPT-6 Astra, first model it rates 'Critical' for cyber capability

    It is the first model OpenAI has released at the 'Critical' top level of its own cyber-risk scale, testing whether deployment safeguards alone can contain that capability.

    Also on When models learned to hack

  40. ReportFrontierUS
    GPT-6 Astra jumps to 62.7% on ARC-AGI-3, six months after models scored 0.5%

    A benchmark designed to resist AI collapsed within months, and the harness-dependent scores show how much results hinge on evaluation setup.

  41. Rule · Enacted, not yet in forceUS-CA
    SB 813 / AB 1405
  42. Rule changeInformationUS-CA
    California enacts SB 1119 requiring child-safety risk assessments for companion chatbots

    Chatbot makers serving Californian minors move from disclosure duties to pre-release risk assessment and audit obligations backed by private lawsuits.

    Also on Children, companions and deepfakes

  43. Rule · Enacted, not yet in forceUS-CA
    SB 1119 / Adam's Law