An open record of AI risk

SafeAI.watch collects public research, reported incidents, warnings, and policy on AI safety and security. Each entry links its source and states the limits of its evidence.

0:00 / 2:27

Transcript
  1. Is AI good or bad for us?

  2. A label squeezes everything a person thinks into one word

    DOOMER

  3. Then people hear the word instead of the person

    ACCELERATIONIST

  4. The same view can end up with opposite labels

    DOOMER

    ACCELERATIONIST

    DECEL

    E/ACC

    LUDDITE

    TECHNO-OPTIMIST

    SAFETYIST

    AI BOOSTER

    EA

    AI BRO

  5. People are people, not camps

  6. AI can read software and find its weak spots

    Mozilla, May 2026

  7. Mozilla used it to fix hundreds of Firefox security flaws in a month

    MOZILLA · FIREFOX SECURITY FIXES

    BENEFITS

    423 fixed in April 2026, 271 of them found by an AI model

  8. Google caught criminals with attack code it believes AI wrote

    GOOGLE · CRIMINAL ATTACK CODE

    REPORTED INCIDENTS

    Google Threat Intelligence Group, May 2026 · its early find may have stopped the attack

  9. AI that describes the world for blind people splits the same way

    SIGHT · AMERICAN FOUNDATION FOR THE BLIND

    American Foundation for the Blind, August 2026

  10. “It can give you a positive infinity of new benefits at the same time that it presents almost a negative infinity of risk”

    Tristan Harris, Center for Humane Technology, July 2026

  11. So how do we keep it on the right side?

  12. We test it

    CONTROLLED TEST

    In a government test, AI agents went onto the live internet and targeted real people

    UK AI Security Institute, August 2026 · internet access left open and safety filters switched off on purpose

  13. “A human maintainer caught and refused to approve the malicious code”

    Same report · no real-world harm found

  14. We fix what tests find

    Government testers probed one lab’s safety monitor and found holes in every version they tried

    SAFEGUARDS

    UK AI Security Institute, July 2026 · each fix was tested again

  15. We make it law

    The largest AI developers in California must report serious safety incidents

    CALIFORNIA · INCIDENT-REPORT LAW

    POLICY & OVERSIGHT

    Transparency in Frontier Artificial Intelligence Act, in effect January 2026 · reports due within 15 days

  16. And we admit what we don’t know

    “no existing study provides a reliable probability of severe loss of control”

    UN Independent International Scientific Panel on AI, September 2026

  17. “clarity creates agency”

    Tristan Harris, TED, April 2025

  18. Use it, value it, and keep your eyes open

  19. Is AI good or bad for us?

    Hold both at once

  20. Stay close to the evidence

    SafeAI.watch

    A public record of AI safety and security

Is AI good or bad for us?

Half of blind and low-vision people who use AI to describe images use it every day (opens in a new tab). One in five say its errors have hurt them. Same tool, same people, same survey. Public arguments about AI rarely hold both facts at once. They sort people into camps with one-word names: doomer, accelerationist, Luddite, techno-optimist, and once a label sticks, people hear the word instead of the person.

Tristan Harris of the Center for Humane Technology, borrowing an image from his co-founder Aza Raskin (opens in a new tab), compares it to looking through one eye at a time: one shows the benefits, the other the risks, and it is hard to open both at once. A label skips the effort. It names the eye someone happened to look through and treats that as the whole person.

The same skill, pointed both ways

AI can now read software and find its weak spots. In April 2026, Mozilla fixed 423 security bugs in Firefox (opens in a new tab), and an AI model found 271 of them. In May, Google's threat intelligence team found criminals holding attack code (opens in a new tab) it believes was built with AI; its early find may have stopped the attack.

This is one capability. A model that finds a flaw fast enough for a defender to fix it can also find one fast enough for an attacker to use it.

Speaking at Davos in 2026 (opens in a new tab), Harris put it in one line: "You can't separate the promise from the peril."

Keeping it on the right side

We use these tools too, and the benefits are real. That is why the work of keeping them safe matters, and much of that work is steady and unglamorous.

People test it. In a 2026 test by the UK AI Security Institute (opens in a new tab), with internet access left open and safety filters switched off on purpose, AI agents went onto the live internet and targeted real people. A human maintainer caught the malicious code before it was approved, and the institute found no real-world harm.

People fix what the tests find. The same institute probed one lab's safety monitor and found holes in every version it tested (opens in a new tab), and each round of fixes was attacked again.

Lawmakers write it into law. Since January 2026, the largest AI developers in California must report serious safety incidents within 15 days (opens in a new tab).

And some questions stay open. A UN scientific panel wrote in September 2026 (opens in a new tab) that "no existing study provides a reliable probability of severe loss of control."

What you can do

Nobody reading this has to solve AI. A more useful question than whether to feel hopeful or afraid is what has to happen next, and who is doing it. Read the source behind a claim. Notice which side you have stopped looking at. Harris's shorthand (opens in a new tab) for why this matters: "clarity creates agency."

SafeAI.watch keeps a dated record of AI safety and security: research, reported incidents, public warnings and policy. Every entry links its original source and separates what happened, what the evidence shows, and what remains uncertain. The aim is to help you hold both at once.